Humberto Luiz Razente

dblp:86/470 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-9687-5730ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Compact Data Structures for the Metric Suffix Array
abstract
The Metric Suffix Array (MSA) is a variant of the classical suffix array designed to support similarity queries over multimedia data. It is a permutation-base index that stores lists of objects ranked by their similarity to a set of reference objects. In this paper, we present two compact alternatives for representing the Metric Suffix Array using less space. The first uses variable-length integer encoding to store the difference between consecutive values in MSA buckets, which are decoded during the queries with a small overhead in time. The second encodes the values in the concatenated ranking list using a Wavelet tree, allowing to obtain MSA values without explicitly computing them. All compact data structures are constructed directly from the input data without the need to compute the complete MSA beforehand, thereby enabling the indexing of much larger data volumes. Experimental results with high-dimensional features extracted from 64 million images showed that our compact alternatives require less than 25% of the original method’s total memory while increasing the query execution time by a small constant factor.
Frederico R. Rosa, Felipe A. Louza, Humberto Luiz Razente
CLEI3
2023 Diversity Similarity Join for Big Data
Yasin N. Silva, Juan Martinez, Pedro Castro Cea, Humberto Luiz Razente, Maria Camila Nardini Barioni
SISAP4
2022 DBSnap-Eval: Identifying Database Query Construction Patterns
abstract
Learning to construct database queries can be a challenging task because students need to learn the specific query language syntax as well as properly understand the effect of each query operator and how multiple operators interact in a query. While some previous studies have looked into the types of database query errors students make and how the availability of expected query results can help to increase the success rate, there is very little that is known regarding the patterns that emerge while students are constructing a query. To be able to look into the process of constructing a query, in this paper we introduce DBSnap-Eval, a tool that supports tree-based queries (similar to SQL query plans) and a block-based querying interface to help separate the syntax and semantics of a query. DBSnap-Eval closely monitors the actions students take to construct a query such as adding a dataset or connecting a dataset with an operator. This paper presents an initial set of results about database query construction patterns using DBSnap-Eval. Particularly, it reports identified patterns in the process students follow to answer common database queries.
Yasin N. Silva, Alexis Loza, Humberto Luiz Razente
ITiCSE (1)3
2022 DBSnap 2: New Features to Construct Database Queries by Snapping Blocks
abstract
Block-based environments for creating computer programs have become very useful learning tools in computer science as they enable focusing on the logic of a program rather than on its syntactical details. While most block-based environments support conventional (imperative) instructions, a few tools have been proposed to create database queries. One of these tools is DBSnap, a highly dynamic and open-source tool to create database query trees by dragging and connecting visual blocks representing datasets and database operators. In this paper, we introduce DBSnap 2, an extension of DBSnap that provides a set of improvements to facilitate the creation of simple and complex queries. The improvements include the support of database views (a key database concept), saving and importing queries, inserting, updating, and deleting data, the creation of charts, and various visual improvements. The demonstration of DBSnap 2 will show how the new features simplify the creation of queries and enable the graphical visualization of query results.
Yasin N. Silva, Alexis Loza, Humberto Luiz Razente
ITiCSE (2)3
2022 Storing data once in M-trees and PM-trees: Revisiting the building principles of metric access methods
Humberto Luiz Razente, Maria Camila Nardini Barioni, Yasin N. Silva
Inf. Syst.1
2021 A comprehensive analysis of delayed insertions in metric access methods
Humberto Luiz Razente, Maria Camila Nardini Barioni, Régis Michel dos Santos Sousa
Inf. Syst.1
2020 EVISClass: a new evaluation method for image data stream classifiers
abstract
Methods for image data stream classification need to update their model constantly and many of these perform this in a supervised way. However, these studies evaluate the performance of their methods assuming that all labels will be available immediately after classification, which is not consistent with various real-world application scenarios. This article proposes a new evaluation method for image data stream classifiers that allows for the exploration of different issues present in real-world applications, such as the emergence of new classes, the evolution of existing classes, and delayed image labels after classification. Through an analysis of the experimental results, we verified that the proposed evaluation method allowed the identification of the issues that most impact the accuracy of the image classifier, indicating a need to direct efforts in carrying out future works to develop strategies to mitigate these issues.
Mateus Curcino de Lima, Maria Camila Nardini Barioni, Elaine Ribeiro de Faria, Humberto Luiz Razente
ICMLA4
2019 Storing Data Once in M-tree and PM-tree
Humberto Luiz Razente, Maria Camila Nardini Barioni
SISAP1
2018 Metric Indexing Assisted by Short-Term Memories
Humberto Luiz Razente, Régis Michel dos Santos Sousa, Maria Camila Nardini Barioni
SISAP1
2011 On query result diversification
abstract
In this paper we describe a general framework for evaluation and optimization of methods for diversifying query results. In these methods, an initial ranking candidate set produced by a query is used to construct a result set, where elements are ranked with respect to relevance and diversity features, i.e., the retrieved elements should be as relevant as possible to the query, and, at the same time, the result set should be as diverse as possible. While addressing relevance is relatively simple and has been heavily studied, diversity is a harder problem to solve. One major contribution of this paper is that, using the above framework, we adapt, implement and evaluate several existing methods for diversifying query results. We also propose two new approaches, namely the Greedy with Marginal Contribution (GMC) and the Greedy Randomized with Neighborhood Expansion (GNE) methods. Another major contribution of this paper is that we present the first thorough experimental evaluation of the various diversification techniques implemented in a common framework. We examine the methods' performance with respect to precision, running time and quality of the result. Our experimental results show that while the proposed methods have higher running times, they achieve precision very close to the optimal, while also providing the best result quality. While GMC is deterministic, the randomized approach (GNE) can achieve better result quality if the user is willing to tradeoff running time.
Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras
ICDE2
2011 Using Visual Analysis to Weight Multiple Signatures to Discriminate Complex Data
abstract
Complex data is usually represented through signatures, which are sets of features describing the data content. Several kinds of complex data allow extracting different signatures from an object, representing complementary data characteristics. However, there is no ground truth of how balancing these signatures to reach an ideal similarity distribution. It depends on the analyst intent, that is, according to the job he/she is performing, a few signatures should have more impact in the data distribution than others. This work presents a new technique, called Visual Signature Weighting (ViSW), which allows interactively analyzing the impact of each signature in the similarity of complex data represented through multiple signatures. Our method provides means to explore the tradeoff of prioritizing signatures over the others, by dynamically changing their weight relation. We also present case studies showing that the technique is useful for global dataset analysis as well as for inspecting subspaces of interest.
Renato Bueno, Daniel S. Kaster, Humberto Luiz Razente, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr.
IV3
2011 DivDB: A System for Diversifying Query Results
Marcos R. Vieira, Humberto Luiz Razente, Maria Camila Nardini Barioni, Marios Hadjieleftheriou, Divesh Srivastava, Caetano Traina Jr., Vassilis J. Tsotras
Proc. VLDB Endow.2
2010 Metric Data Analysis Enhanced through Temporal Visualization
abstract
The human vision can naturally interpret data in spaces of 2 or 3 dimensions. When data is in higher dimensional spaces, in most cases the visualization is not intuitive. Regarding metric spaces, the interpretation is even harder, since they often do not have a direct spatial representation. However, the need to analyze how metric-represented data evolve over time is pretty common when one needs to understand several phenomena and in decision making processes, as it occurs in medical and agrometeorological applications. This paper presents three interactive techniques to visualize metric data that vary over time. Each one focus on a different way to interpret the temporal information. The first technique shows data evolving in a timeline axis. The second overlaps evolving snapshots of the space showing how the space varies regarding time. The last one does not treat temporal data as a dimension, it is used instead to define the similarity among complex data, employing the new concept of metric-temporal spaces, which seamlessly integrate time and metric data into a single similarity space. Visualization examples with real datasets are presented to show the usefulness of the proposed techniques.
Renato Bueno, Humberto Luiz Razente, Daniel S. Kaster, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr.
IV2
2009 Seamlessly integrating similarity queries in SQL
abstract
Abstract Modern database applications are increasingly employing database management systems (DBMS) to store multimedia and other complex data. To adequately support the queries required to retrieve these kinds of data, the DBMS need to answer similarity queries. However, the standard structured query language (SQL) does not provide effective support for such queries. This paper proposes an extension to SQL that seamlessly integrates syntactical constructions to express similarity predicates to the existing SQL syntax and describes the implementation of a similarity retrieval engine that allows posing similarity queries using the language extension in a relational DBMS. The engine allows the evaluation of every aspect of the proposed extension, including the data definition language and data manipulation language statements, and employs metric access methods to accelerate the queries. Copyright © 2008 John Wiley & Sons, Ltd.
Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr.
Softw. Pract. Exp.2
2008 A novel optimization approach to efficiently process aggregate similarity queries in metric access methods
abstract
A similarity query considers an element as the query center and searches a dataset to find either the elements far up to a bounding radius or the k nearest ones from the query center. Several algorithms have been developed to efficiently execute similarity queries. However, there are queries that require more than one center, which we call Aggregate Similarity Queries. Such queries appear when the user gives multiple desirable examples, and requests data elements that are similar to all of the examples, as in the case of applying relevance feedback. Here we give the first algorithms that can handle aggregate similarity queries on Metric Access Methods (MAM) such as the M-tree and Slim-tree. Our method, which we call Metric Aggregate Similarity Search (MASS) has the following properties: (a) it requires only the triangle inequality property; (b) it guarantees no false-dismissals, as we prove that it lower-bounds the aggregate distance scores; (c) it can work with any MAM; (d) it can handle any number of query centers, which are either scattered all over the space or concentrated on a restricted region. Experiments on both real and synthetic data show that our method scales on both the number of elements and, if the dataset is in a spatial domain, also on its dimensionality. Moreover, it achieves better results than previous related methods.
Humberto Luiz Razente, Maria Camila Nardini Barioni, Agma J. M. Traina, Christos Faloutsos, Caetano Traina Jr.
CIKM1
2008 Accelerating k-medoid-based algorithms through metric access methods
Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr.
J. Syst. Softw.2
2006 SIREN: A Similarity Retrieval Engine for Complex Data
Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr.
VLDB2
2002 Extending Relational atabases to Support Content-based Retrieval of Medical Images
abstract
This paper shows how to support images in a relational database, so it can fulfill the requirements to be used as the storage mechanism of a PACS. This support includes the ability to answer similarity queries based on the image content, providing fast image retrieval based on indexing structures. The main concept allowing this support is the definition of distance functions based on features, which are extracted from the images as they are stored in the database. An extension to SQL enables the construction of an interpreter that intercepts the extended commands and translates them into standard SQL, allowing one to take advantage of any relational database server. We describe experiments made with a prototype implemented using these concepts, which allowed answering queries up to 20 times faster than using existing relational servers alone.
Myrian R. B. Araujo, Caetano Traina Jr., Agma J. M. Traina, Josiane Maria Bueno, Humberto Luiz Razente
CBMS5