VLDB 2026 Research / reviewers in the wild / expert
Agma J. M. Traina
dblp:t/AgmaJMTraina · also Agma Juci Machado Traina
· DBLP profile ↗
162ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0003-4929-7258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 91 · 6 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 70 · 6 first-author · 17 since 2021Databases, data management, data science and information retrieval · 64 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 64 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A review on Python libraries for temporal network analysisabstractContext: Complex networks represent systems with non-trivial connections and are widely used in fields such as social media, biology, and transportation. Temporal networks extend this by capturing the evolution of connections over time, providing insights into event sequences and information diffusion. Analyzing these networks requires specialized tools, and Python offers a variety of libraries tailored for this purpose. Objective: This study evaluates Python libraries designed for temporal network analysis based on multiple criteria. The aim is to assess the strengths and limitations of these tools, guide users in selecting appropriate libraries, and identify gaps for future development. Methods: A comparative analysis was conducted on selected Python libraries using predefined evaluation criteria. The assessment considered factors such as available documentation, supported metrics, visualization capabilities, supported format, uniqueness, community support, and popularity. Data were gathered from official documentation, community forums, scientific papers, and usage statistics. Results: Findings indicate that the TGX, Teneto, and PathpyG stand out, excelling in three of five criteria. Networkx-t shows balanced performance with no significant drawbacks, making it a reliable general-purpose choice. However, several tools have limitations in specific areas, such as a lack of comprehensive documentation or advanced visualization features. Conclusion: This review provides an overview of existing Python tools for temporal network analysis, offering insights into their capabilities and shortcomings. The results assist researchers and practitioners in selecting suitable libraries while highlighting areas for improvement and potential future developments in the field. Claudio D. G. Linhares, Jean R. Ponciano, Martim R. Oliveira, Amílcar Soares Júnior 0001, Agma J. M. Traina, Andreas Kerren |
Inf. Softw. Technol. | 5 |
| 2025 | FairMed-FL: Federated Learning for Fair and Unbiased Deep Learning in Medical ImagingabstractDeep learning models have achieved great success in medical imaging tasks. However, recent work on fairness in healthcare has shown these models can be biased, potentially leading to discriminatory treatment of patients based on demographic attributes such as race, gender, and age. Data bias, often resulting from imbalanced and non-representative datasets, can negatively impact model fairness. While aggregating data from multiple sources can help mitigate data bias, privacy concerns make the data-sharing process challenging. In this scenario, Federated Learning (FL) has emerged as a solution for the collaborative training of models without data sharing. This paper presents FairMed-FL, a methodology to assess fairness in FL for medical imaging tasks. By utilizing two public chest X-ray datasets partitioned by sex and age, we compared federated models trained with clients from single and multiple datasets against centralized models trained on each client's data. The results indicate that FL reduces performance discrepancies between demographic groups, enhances the performance of the worstperforming groups, and improves overall metrics compared to centralized approaches. These findings highlight its potential for promoting fairness in medical imaging. Êrica Peters do Carmo, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 3 |
| 2025 | Improving U-Net with Attention Mechanism for Medical Image Segmentation ApplicationsabstractMedical image segmentation plays a vital role in numerous applications and has gained significant attention since the introduction of the U-Net model, which enabled convolutional neural networks to achieve high performance with manageable computational costs. Recently, attention mechanisms have emerged as a promising approach to enhance model performance by emphasizing relevant features while suppressing irrelevant ones. This study explores the integration of channel and spatial attention mechanisms into the U-Net architecture, evaluating their impact on segmentation performance and computational cost. Experiments conducted on six public medical imaging datasets demonstrated performance improvements, with Intersection over Union (IoU) gains ranging from 1.62% to 33.66% compared to the original U-Net. These results highlight the potential of attention mechanisms to significantly improve the efficiency and effectiveness of medical image segmentation models. Eduardo Sperle Honorato, Mariana Aya Suzuki Uchida, Agma J. M. Traina, Denis F. Wolf |
CBMS | 3 |
| 2025 | Modeling Clinical Data with Attention: A Knowledge Graph Approach with CliniKGabstractGiven a set of historical, unlabeled patient records, how can we model the relationships between various concepts related to health conditions and treatments? Modeling health data as knowledge graphs can aid in better understanding and mining recurrent and abnormal patterns within massive amounts of records. However, for generic data modeling, the concepts and relationships of the data are often defined manually, which can be laborious and prone to human errors. The lack of structure in textual reports makes modeling challenging, as there is typically no standard in information, terminology, or other elements. This work proposes CliniKG for utilizing Large Language Models (LLMs) to structure patient records, enabling the modeling of health data concepts. First, we employ LLMs with zero-, or few-shot learning to define the graph's relationships of pairwise concepts. Based on the resulting modeling, we extract meaningful features from nodes and define their semantics automatically. With data visualizations, CliniKG highlights key findings, such as recurrent or rare relationships. The experimental evaluation shows CliniKG in action using a real dataset from a public hospital in Brazil. We performed a qualitative analysis with 40 domain experts to evaluate medium-sized LLMs and prompt con-figurations. The study reveals interesting patterns automatically identified within the data. CliniKG exhibits linear performance relative to the number of nodes, and the visual tools can assist specialists in monitoring patients' conditions. Eduardo Moura, Rafael C. G. Conrado, Leonardo de Oliveira Campos, Mauro M. Olivatto, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina, Mirela Teixeira Cazzolato |
CBMS | 7 |
| 2025 | AI-Driven Public Health Surveillance: Analyzing Vulnerable Areas in Brazil Using Remote Sensing and Socioeconomic DataabstractUrban vulnerability assessment is crucial for understanding the spatial distribution of deprived areas and associated risks. Slum residents face significantly worse health outcomes than non-slum urban populations, with neighborhood effects being critical in social epidemiology. Identifying such areas is vital because they present public health challenges that climate change and increased air pollution can exacerbate. Accordingly, this study proposed an AI-driven methodology that integrates remote sensing data, socioeconomic indicators, and machine learning algorithms to identify and analyze vulnerable areas in Brazil. To create a vulnerability index, we incorporate multiple data sources, including Sentinel-2 and Sentinel-5P imagery, Brazilian socioeconomic indicators, and OpenStreetMap. Hence, we predicted pollution indicators using regression algorithms such as Random Forest, XGBoost, and Linear Regression. Our findings demonstrate that integrating multi-source data is a promising approach for better understanding deprived areas, indicating that slums (called “favelas” in Brazil) exhibit an intense concentration of the sociocconomic vulnerability index, a key determinant of deprivation. However, non-slum areas may present heterogeneous conditions, with some regions showing vulnerability levels comparable to those of slums while others show better conditions. Our results highlight the potential of AIdriven approaches for urban vulnerability assessment, offering insights for policymakers and researchers. Joao Pedro Silva, Erikson Júlio De Aguiar, Gabriel Spadon, Agma J. M. Traina, Jose F. Rodrigues |
CBMS | 4 |
| 2025 | Cosim-Gres: Towards Similarity Queries Optimization Inside RDBMSabstractABSTRACT Introduction This paper presents CoSIM‐Gres, a new module implemented over Postgres capable of performing exact similarity searches to answer both Range and ‐ queries, using any of three access methods: Sequential Access, the Slim‐tree Metric Access Method (MAM), or the Gist R‐tree. To the best of the authors' knowledge, this is the first system currently capable of performing both types of queries with the possibility of choosing different access methods while maintaining full integration of the similarity‐related syntax with the other SQL statements. Contribution Our main contribution is an in‐depth comparison of cases when each access method is better for processing similarity queries. It is an essential first step that any attempt to optimize similarity queries within DBMS must consider. Methods Experiments were performed to compare Slim‐tree with Sequential Access and Gist R‐tree available in Postgres Cube extension, analyzing the impact of varying dimensionality and distance functions on the execution time. Results They show that the Slim‐tree is up to 18.0 times faster than Sequential Access, whereas the Gist R‐tree may be up to 5.8 times faster than the Slim‐tree, although the Gist R‐tree in Postgres is restricted to index only dimensional data, with at most 100 dimensions, and is applicable only for ‐ queries. The experiments also revealed that the growth of data dimensionality negatively impacts both the MAM and Gist R‐tree performance when compared to the Sequential Access. Also, more complex distance functions reduce the advantage of MAM over Sequential Access. This makes choosing the best option to execute a query a decision that should be carefully evaluated before the query execution. Igor A. R. Eleutério, Willian D. Oliveira, Larissa Roberta Teixeira, Thiago Galbiatti Vespa, William Zaniboni Silva, Agma J. M. Traina, Caetano Traina Jr. |
Softw. Pract. Exp. | 6 |
| 2025 | ZigzagNetVis: Suggesting Temporal Resolutions for Graph Visualization Using Zigzag PersistenceabstractTemporal graphs are commonly used to represent complex systems and track the evolution of their constituents over time. Visualizing these graphs is crucial as it allows one to quickly identify anomalies, trends, patterns, and other properties that facilitate better decision-making. In this context, selecting an appropriate temporal resolution is essential for constructing and visually analyzing the layout. The choice of resolution is particularly important, especially when dealing with temporally sparse graphs. In such cases, changing the temporal resolution by grouping events (i.e., edges) from consecutive timestamps - a technique known as timeslicing - can aid in the analysis and reveal patterns that might not be discernible otherwise. However, selecting an appropriate temporal resolution is a challenging task. In this paper, we propose ZigzagNetVis, a methodology that suggests temporal resolutions potentially relevant for analyzing a given graph, i.e., resolutions that lead to substantial topological changes in the graph structure. ZigzagNetVis achieves this by leveraging zigzag persistent homology, a well-established technique from Topological Data Analysis (TDA). To improve visual graph analysis, ZigzagNetVis incorporates the colored barcode, a novel timeline-based visualization inspired by persistence barcodes commonly used in TDA. We also contribute with a web-based system prototype that implements suggestion methodology and visualization tools. Finally, we demonstrate the usefulness and effectiveness of ZigzagNetVis through a usage scenario, a user study with 27 participants, and a detailed quantitative evaluation. Raphaël Tinarrage, Jean R. Ponciano, Claudio D. G. Linhares, Agma J. M. Traina, Jorge Poco |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | MedTimeSplit: Continual dataset partitioning to mimic real-world settings for federated learning on Non-IID medical image dataabstractTraditional Deep Learning (DL) approaches for medical image classification rely on centralized, static datasets, which do not adequately reflect the dynamic, real-world medical practice where data is continually generated. In contrast, Federated Learning (FL) enables decentralized model training on localized data while preserving privacy. Yet, current FL methods struggle to handle Non-Identically Independently Distributed (Non-IID) data streams over time. This paper introduces Med-TimeSplit, a novel dataset partitioning strategy that integrates Online Continual Learning (OCL) with FL to simulate real-world medical data flows more realistically and effectively. MedTimeSplit partitions data into Non-IID, time-based increments, mimicking dynamic sourcing in medical environments. We evaluate its impact on FL model performance for medical image classification, focusing on skin lesions, and analyze the system’s resilience to backdoor attacks. Our experiments demonstrate that MedTimeSplit outperforms existing methods in both accuracy and robustness, offering a viable solution for real-world medical applications. Additionally, we propose new metrics to measure model behavior over time, including Average Bad Decisions (ABD) and Overall Changing Mistakes (OCM), which provide deeper insights into model performance specifically under OCL conditions. The results highlight the promise of combining OCL with FL in the medical domain, paving the way for more secure and adaptive healthcare solutions. The source code is available on GitHub1. Erikson Júlio De Aguiar, Agma J. M. Traina, Abdelsalam Helal |
IEEE Big Data | 2 |
| 2024 | RADAR-MIX: How to Uncover Adversarial Attacks in Medical Image Analysis through ExplainabilityabstractMedical image analysis is an important asset in the clinical process, providing resources to assist physicians in detecting diseases and making accurate diagnoses. Deep Learning (DL) models have been widely applied in these tasks, improving the ability to recognize patterns, including accurate and fast diagnosis. However, DL can present issues related to security violations that reduce the system’s confidence. Uncovering these attacks before they happen and visualizing their behavior is challenging. Current solutions are limited to binary analysis of the problem, only classifying the sample into attacked or not attacked. In this paper, we propose the RADAR-MIX framework for uncovering adversarial attacks using quantitative metrics and analysis of the attack’s behavior based on visual analysis. The RADAR-MIX provides a framework to assist practitioners in checking the possibility of adversarial examples in medical applications. Our experimental evaluation shows that the Deep-Fool and Carlini & Wagner (CW) attacks significantly evade the ResNet50V2 with a slight noise level of 0.001. Furthermore, our results revealed that the gradient-based methods, such as Gradient-weighted Class Activation Mapping (Grad-CAM) and SHapley Additive exPlanations (SHAP), achieved high attack detection effectiveness. While Local Interpretable Model-agnostic Explanations (LIME) presents low consistency, implying the most ability to uncover robust attacks supported by visual analysis. Erikson Júlio De Aguiar, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 3 |
| 2024 | A symptom-based community-weighted similarity approach for inpatient health condition monitoringabstractGiven a patient’s series of exams conducted over time, how can we identify cases with similar abnormalities or symptoms? Hospitals and medical facilities continuously monitor patients through periodic exams, a crucial practice for assessing their current condition and potential progression, thereby supporting decision-making. However, similarity-based searches often consider several exams of a patient, most times overlooking the temporal aspect, which is crucial for patient monitoring. In this paper, we present: (1) a novel similarity search framework that identifies similar cases based on symptoms while considering the temporal evolution of the patients’ conditions; and (2) a novel similarity function, called GCWei function, which is built upon the traditional Levenshtein similarity and improves the quality of the search by penalizing the similarity between non-related sets of symptoms. To identify relations, GCWei relies on well-established graph community detection procedures using all patients’ historical data. By combining (1) and (2), we obtain a search approach called GCWei-based search, which efficiently retrieves similar cases with similar developments and thus gives the specialist a broader view of the patient’s condition based on past cases of other patients. To demonstrate the value of our approach, we evaluate it both quantitatively and qualitatively using the recent and publicly available MIMIC-IV database. Jean R. Ponciano, Mirela Teixeira Cazzolato, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 5 |
| 2023 | TgrApp: Anomaly Detection and Visualization of Large-Scale Call GraphsabstractGiven a million-scale dataset of who-calls-whom data containing imperfect labels, how can we detect existing and new fraud patterns? We propose TgrApp, which extracts carefully designed features and provides visualizations to assist analysts in spotting fraudsters and suspicious behavior. Our TgrApp method has the following properties: (a) Scalable, as it is linear on the input size; and (b) Effective, as it allows natural interaction with human analysts, and is applicable in both supervised and unsupervised settings. Mirela Teixeira Cazzolato, Saranya Vijayakumar, Namyong Park 0001, Meng-Chieh Lee, Polo Chau, Pedro Fidalgo, Bruno Lages, Agma J. M. Traina, Christos Faloutsos |
AAAI | 9 |
| 2023 | Assessing Vulnerabilities of Deep Learning Explainability in Medical Image Analysis Under Adversarial SettingsabstractDeep Learning (DL) is a valuable set of techniques that improve medical decision-making based on imaging exams, such as Chest X-rays (CXR), Computed Tomography (CT), and Optical Coherence Tomography (OCT). However, DL models may be susceptible to adversarial attacks when perturbed (tam-pered) examples sneak into the data, decreasing the model's confidence. In this paper, we evaluate the vulnerabilities of DL applied to medical images and analyze the effects of attacks on the Gradient-weighted Class Activation Mapping (GRAD-CAM). Our experiments were conducted on two scenarios: (i) CXR images with binary class; (ii) OCT images with multi-class. Vulnerabilities are described by Fooling Rate (FR) and visual analysis of Grad-CAM. We show that the PGD is the most malicious deed for multi-class, reaching an FR of up to 96%, whereas DeepFool is hurtful for binary classes, reaching an FR of up to 93%. Our analysis can be used to understand the adversarial attacks over medical images and their effects on explainability. The developed code is available at GitHub11https://github.com/eriksonJAguiar/Grad-Attacks-CBMS-2023. Erikson Júlio De Aguiar, Márcus V. L. Costa, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 4 |
| 2023 | Exploratory Data Analysis in Electronic Health Records Graphs: Intuitive Features and Visualization ToolsabstractGiven a large, unlabeled set of Electronic Health Records (EHRs) acquired from multiple hospitals, how can we analyze the available entities and identify relationships in the data? Also, how can we perform Exploratory Data Analysis (EDA) over such EHR data? Many medical institutions generate EHRs as tabular data with entities and attributes in common. However, due to a large number of records, attributes, and high cardinality, exploring the different datasets and finding patterns and insights become laborious and prone to errors. In this work, we propose GraF- Eda for EDA over EHR data from different institutions. GraF-EDA models EHRs as time-evolving graphs, allowing the interoperability of such data into a single representation. We extract meaningful features from the graph nodes and provide intuitive visualizations to improve data explainability. We evaluate GraF-EDA with four COVID-19 datasets from hospitals of the São Paulo state, Brazil, resulting in million-scale graphs. Our method identified correlations, similarities and dissimilarities among medical treatments, exams, clinics, and outcomes. With the visual tools provided by GraF-EDA, we were able to spot cases of interest and check more details about them. Our results indicate that GraF-EDA is a fast, effective, open-sourced tool for EDA of EHRs from multiple institutions. Mirela Teixeira Cazzolato, Marco A. Gutierrez 0001, Caetano Traina Jr., Christos Faloutsos, Agma J. M. Traina |
CBMS | 5 |
| 2023 | A Deep Learning-based Radiomics Approach for COVID-19 Detection from CXR Images using Ensemble Learning ModelabstractMedical image analysis plays a major role in aiding physicians in decision-making. Specifically in detecting COVID-19, Deep Learning (DL) and radiomic approaches have achieved promising results separately. However, DL results are hard to interpret/visualize, and the radiomic approach encompasses successive steps, such as image acquisition, image processing, segmentation, feature extraction, and analysis. In this paper, we integrate DL with radiomic approaches, aiding in detecting COVID-19. We use DL models to extract 128 relevant deep radiomic features to assess COVID-19 from several image sources of 392 representative chest X-ray (CXR) exams. We avoid successive radiomic steps by employing DL (transfer learning) from Imagenet's VGG-16, ResNet50V2, and DenseNet201 networks. We considered a set of Machine Learning (ML) algorithms to further validate our results, providing an ensemble model to detect COVID-19. Our experimental results show that our approach achieved 95% AUC using 128 relevant features from DenseNet201. Conversely, our ensemble model presented 91% AUC, indicating that deep learning-based radiomics could increase binary classification performance in a real scenario. In addition, we highlight that our approach can be adapted to create other DL-based radiomics tools. For reproducibility, we made our code available at https://github.com/usmarcv/CBMS-DL-based-radiomics. Márcus V. L. Costa, Erikson Júlio De Aguiar, Lucas Santiago Rodrigues, Jonathan S. Ramos, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 6 |
| 2023 | CallMine: Fraud Detection and Visualization of Million-Scale Call GraphsabstractGiven a million-scale dataset of who-calls-whom data containing imperfect labels, how can we detect existing and new fraud patterns? We propose CallMine, with carefully designed features and visualizations. Our CallMine method has the following properties: (a) Scalable, being linear on the input size, handling about 35 million records in around one hour on a stock laptop; (b) Effective, allowing natural interaction with human analysts; (c) Flexible, being applicable in both supervised and unsupervised settings; (d) Automatic, requiring no user-defined parameters. Mirela Teixeira Cazzolato, Saranya Vijayakumar, Meng-Chieh Lee, Catalina Vajiac, Namyong Park 0001, Pedro Fidalgo, Agma J. M. Traina, Christos Faloutsos |
CIKM | 7 |
| 2023 | Pushing diversity into higher dimensions: The LID effect on diversified similarity searching
Daniel L. Jasbick, Lúcio F. D. Santos, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Daniel de Oliveira 0001, Marcos V. N. Bedo |
Inf. Syst. | 4 |
| 2023 | ClinicalPath: A Visualization Tool to Improve the Evaluation of Electronic Health Records in Clinical Decision-MakingabstractPhysicians work at a very tight schedule and need decision-making support tools to help on improving and doing their work in a timely and dependable manner. Examining piles of sheets with test results and using systems with little visualization support to provide diagnostics is daunting, but that is still the usual way for the physicians' daily procedure, especially in developing countries. Electronic Health Records systems have been designed to keep the patients' history and reduce the time spent analyzing the patient's data. However, better tools to support decision-making are still needed. In this article, we propose ClinicalPath, a visualization tool for users to track a patient's clinical path through a series of tests and data, which can aid in treatments and diagnoses. Our proposal is focused on patient's data analysis, presenting the test results and clinical history longitudinally. Both the visualization design and the system functionality were developed in close collaboration with experts in the medical domain to ensure a right fit of the technical solutions and the real needs of the professionals. We validated the proposed visualization based on case studies and user assessments through tasks based on the physician's daily activities. Our results show that our proposed system improves the physicians' experience in decision-making tasks, made with more confidence and better usage of the physicians' time, allowing them to take other needed care for the patients. Claudio D. G. Linhares, Daniel Mario de Lima, Jean R. Ponciano, Mauro M. Olivatto, Marco A. Gutierrez 0001, Jorge Poco, Caetano Traina Jr., Agma J. M. Traina |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | LargeNetVis: Visual Exploration of Large Temporal Networks Based on Community TaxonomiesabstractTemporal (or time-evolving) networks are commonly used to model complex systems and the evolution of their components throughout time. Although these networks can be analyzed by different means, visual analytics stands out as an effective way for a pre-analysis before doing quantitative/statistical analyses to identify patterns, anomalies, and other behaviors in the data, thus leading to new insights and better decision-making. However, the large number of nodes, edges, and/or timestamps in many real-world networks may lead to polluted layouts that make the analysis inefficient or even infeasible. In this paper, we propose LargeNetVis, a web-based visual analytics system designed to assist in analyzing small and large temporal networks. It successfully achieves this goal by leveraging three taxonomies focused on network communities to guide the visual exploration process. The system is composed of four interactive visual components: the first (Taxonomy Matrix) presents a summary of the network characteristics, the second (Global View) gives an overview of the network evolution, the third (a node-link diagram) enables community- and node-level structural analysis, and the fourth (a Temporal Activity Map - TAM) shows the community- and node-level activity under a temporal perspective. We demonstrate the usefulness and effectiveness of LargeNetVis through two usage scenarios and a user study with 14 participants. Claudio D. G. Linhares, Jean R. Ponciano, Diogenes S. Pedro, Luis Enrique Correa da Rocha, Agma J. M. Traina, Jorge Poco |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | ORTree: Tuning Diversified Similarity Queries by Means of Data Partitioning
João V. O. Novaes, Lúcio F. D. Santos, Agma J. M. Traina, Caetano Traina Jr. |
ADBIS | 3 |
| 2022 | TgraphSpot: Fast and Effective Anomaly Detection for Time-Evolving GraphsabstractGiven a large, time-evolving graph of who-calls-whom-when, how can we help analysts find anomalies and fraudsters? How can we explain our decisions? We provide TgraphSpot, which carefully extracts features that are often related to fraud; and which provides informative, interactive plots that help analysts zoom down to the few strange nodes. We present the architecture and design decisions of TgraphSpot. Thanks to our careful feature-extraction algorithms, it scales linearly, taking 2.5 hours on a stock laptop, to process 29 million phone calls. More importantly, when applied on a real dataset of millions of phone calls, it discovered suspicious nodes; experts confirmed that those nodes are fraudsters that had been undetected so far. Mirela Teixeira Cazzolato, Saranya Vijayakumar, Namyong Park 0001, Meng-Chieh Lee, Pedro Fidalgo, Bruno Lages, Agma J. M. Traina, Christos Faloutsos |
IEEE Big Data | 8 |
| 2022 | Wia-Spine: A CBIR environment with embedded radiomic features to assess fragility fracturesabstractOsteoporosis is a systemic disorder that reduces the bone mineral density, increasing the vertebrae's fragility and proneness to fracture. Although the bone densitometry index t-Score is a solid marker for the osteoporosis diagnosis, its measure alone is insufficient to predict the future development of fragility fractures. A complementary approach to address vertebral bone characterization is the analysis of magnetic resonance imaging (MRI) by radiomic features, which model vertebral bodies' morphological properties after color and texture. Radiomic features have been employed for detecting fragility fractures in related work, but, to the best of our knowledge, no study has been conducted on their suitability to recover similar, diagnosed cases that could hint at future fractures. We fulfill this gap by designing a Content-based Image Retrieval (CBIR) tool with embedded radiomic features, which uses past cases recovered from an annotated database to (i) identify an existing fragility fracture in a query vertebra and (ii) predict a fracture to a query vertebra from an aging patient. The proposed CBIR was evaluated on a reference database of 273 vertebral bodies from sagittal T2-weighted MRIs. The results indicate our fine-tuned approach spotted fragility fractures accurately$(\mathrm{F}1-\text{Score} =0.83,\ \text{Precision} =0.83,\ \text{AUC} =0.81,\ \text{CI} =95\%)$. We also investigated the CBIR potential to predict fractures in a case study regarding three patients from the reference database (confirmed osteoporosis, MRI in [2012–2017]). The system correctly inferred the prediction of future fractures for query vertebrae, which were confirmed a few years later (MRI in [2018–2021]). Such empirical findings suggest CBIR can support a differential diagnosis in the assessment of local fragility fractures. Marcos V. N. Bedo, Jonathan S. Ramos, Agma J. M. Traina, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Paulo Mazzoncini de Azevedo Marques |
CBMS | 3 |
| 2022 | Analysis of vertebrae without fracture on spine MRI to assess bone fragility: A Comparison of Traditional Machine Learning and Deep LearningabstractBone mineral density (BMD) is the international standard for evaluating osteoporosis/osteopenia. The success rate of BMD alone in estimating the risk of vertebral fragility fracture (VFF) is approximately 50%, making BMD far from ideal in predicting VFF. In addition, whether or not a patient has been diagnosed with osteoporosis or osteopenia, he or she may suffer a VFF. For this reason, we conducted an extensive empirical study to assess VFFs in postmenopausal women. We considered a representative dataset of 94 T1- and T2-weighted routine spine MRI (with osteopenia or osteoporosis), split into 2,400 samples (slices). Comparing the classification results of machine learning and deep learning (DL) techniques showed that DL generally achieved better results at the cost of higher computational power and hard explainability. ResNet achieved the best results in discriminating patients from groups with and without VFFs with 83% accuracy and 90% AUC (with a confidence interval of 99%). Our results represent a significant step toward prospective and longitudinal studies investigating methods to achieve higher accuracy in predicting VFFs based on spine MRI features of vertebrae without fracture. Jonathan S. Ramos, Erikson Júlio De Aguiar, Ivar Vargas Belizario, Márcus V. L. Costa, Jamilly G. Maciel, Mirela Teixeira Cazzolato, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Agma J. M. Traina |
CBMS | 9 |
| 2022 | Variational Autoencoders for Medical Image RetrievalabstractThis paper presents an approach based on Variational Autoencoders (VAEs) for unsupervised learning of deep features for Content-Based Medical Image Retrieval (CBMIR). We show that this unsupervised approach can yield better results than the predominant supervised approach based on classification, and can even be used in combination with the classification approach, resulting in a second, mixed model. Despite the VAEs retrieving images more visually similar, the evaluation methodology usually employed in literature is not able to reveal this advantage, which is important for CBMIR. We then propose a new evaluation method based on hidden classes and show that it reflects the visual similarity of the retrieved images better than the traditional evaluation. Cézanne Alves, Agma J. M. Traina |
INISTA | 2 |
| 2022 | Establishing trajectories of moving objects without identities: The intricacies of cell tracking and a solutionabstractStoring, querying, predicting, and interpolating trajectories of moving objects is a topic which the database community has studied for decades. We study a new variant of this problem in this article: We deal with a set of moving objects which do not have an identity, i.e., one does not know whether an object is identical to one observed earlier at another position. Our use case is a stream of images of cells of developing embryos. There exist so-called tracking tools. They match cells in such image sequences, to build trajectory vectors. However, these trackers have certain weaknesses, including counter-intuitive parameters and the expectation of users manually correcting trajectories. In this paper, we propose fully automatic tracking algorithms. They rely on space partitioning heuristics to match cells. This gives way to much cheaper data-analysis pipelines, as we will explain. We also propose two algorithms predicting the next positions of cells, given earlier ones. Experiments over 12 datasets show that our new approaches reduce the execution time by up to 7.8 times for tracking and 6.2 times for prediction. Prediction quality increases by up to 5.6% over the best tracker. • Cells can be modeled as moving objects without identity that move under uncertainty. • Predictors establish cell motion accurately based on observed cell positions. • Cell prediction avoids computationally costly steps of the tracking pipeline. Mirela Teixeira Cazzolato, Agma J. M. Traina, Klemens Böhm |
Inf. Syst. | 2 |
| 2021 | Multilevel Clustering Explainer: An Explainable Approach to Electronic Health RecordsabstractMachine learning (ML) algorithms have been used in many areas of activity, and their results can often be applied without further human intervention. The ML algorithms have also been widely used in medical contexts, but in this area, the result needs to be thoroughly confirmed by a specialist, who needs explanatory information on how the results were obtained. Aimed at such scenarios, we propose the Multilevel Clustering Explainer (MCE), a method capable of providing explanatory information to health professionals about the knowledge discovery process. The MCE was developed for the analysis of medical data, providing a synthesis of explanatory information for the specialist to quickly and clearly understand how the results were obtained. José Maria Clementino, Bruno S. Faiçal, Christian C. Bones, Caetano Traina Jr., Marco A. Gutierrez 0001, Agma J. M. Traina |
CBMS | 6 |
| 2021 | I-CovidVis - A Visual Analytics Tool for Interoperable Healthcare Databases using GraphsabstractThe current COVID-19 pandemic has promoted the periodic release of several health databases aimed at discovering relationships in the data, detecting similar problems in patients, and studying the evolution of the disease. A way to exploit the data is to use visualization techniques, which can lead to the discovery of insights and patterns, as well as to guide analysis procedures to understand the data. In this paper, we present I-CovidVis, a visualization tool to explore data from interoperable healthcare systems, able to compare and navigate in both global and local perspectives. Our approach is to model data as a graph and explore its structural and temporal views. Our proposal facilitates the perception of patterns, trends, periodicity, and anomalies, resulting in faster decision making. Claudio D. G. Linhares, Daniel Mario de Lima, Christian C. Bones, Marina de Sá Rebelo, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 7 |
| 2021 | BEAUT: a radiomic approach to identify potential lumbar fractures in magnetic resonance imagingabstractBone densitometry (DEXA) is the international reference standard to evaluate Bone Mineral Density (BMD) and diagnose osteoporosis. However, DEXA is far from ideal when used to predict fragility fractures, which are strongly related to morbidity and mortality. According to the literature, spine MRI texture features correlate well with DEXA measurements. For this reason, we conducted an extensive empirical study aimed at assessing fragility fractures secondary to osteoporosis. To perform the evaluations, we developed a radiomic-based approach called BEAUT (BonE Analysis Using Texture). We performed experiments on a meaningful database composed of 47 T2-weighted sagittal sequences from lumbar spine MRI. The patients were diagnosed with osteopenia or osteoporosis according to DEXA (patients with low bone mass). BEAUT achieved an accuracy of 92% and 97% AUC with feature selection to discriminate between patients from groups `Fractures' and `No Fractures'. The results support claiming that texture features potentially discriminate subjects with bone mass loss, spotting those at risk of fragility fractures. Jonathan S. Ramos, Jamilly G. Maciel, Mirela Teixeira Cazzolato, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Agma J. M. Traina |
CBMS | 6 |
| 2020 | Semi-Automatic Ulcer Segmentation and Wound Area Measurement Supporting TelemedicineabstractMany patients suffer from chronic skin lesions, commonly known as ulcers. The size evolution of chronic wounds provides meaningful clues regarding the patient's clinical state for healthcare professionals and caretakers. Many studies have been proposed in recent years to support the treatment of skin ulcers. However, there is a lack of practical solutions, as existing studies are not targeted at immediate use in daily medical practice. In this work, we propose URule, an essentially practical framework for segmentation and measurement of skin ulcers. URule-App, a mobile instance of the framework, analyzes images taken by a common camera from a mobile device. The segmentation requires the user to manually outline the outsider region of both the wound and the measurement tool. URule-Seg segments the image and estimates the wound area. The user can further improve the estimated area by manually informing the span of a centimeter in the image. The experimental evaluation reveals that URule can accurately segment ulcer wounds semi-automatically, with an average F-Measure of 0.8 for segmentation, and processing measurement tools better than the manual process in three out of five tested rulers. Mirela Teixeira Cazzolato, Jonathan S. Ramos, Lucas Santiago Rodrigues, Lucas C. Scabora, Daniel Y. T. Chino, Ana Elisa Serafim Jorge, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 9 |
| 2020 | Bag-of-Attributes Representation: A Vector Space Model for Electronic Health Records Analysis in OMOPabstractSeveral studies have been performed worldwide to improve health services using data generated by digital medical systems. The increasing volume of data generated by these systems is making the use of knowledge discovery and data analysis techniques essential to improve the quality of the health services, which are offered by the medical facilities. However, it is possible to observe a gap, in the literature, about generic and flexible vector space models (VSM) that are well adapted to handle electronic health records (EHR), requiring that each knowledge discovery effort develop their own VSM or other representation model. This restriction can turn a knowledge discovery task over clinical pathways nonviable for comparative evaluations among different methods. Targeting such scenario, we propose the Bag-of-Attributes Representation (BOAR). BOAR represents an EHR as an n-dimensional vector space. Since BOAR takes advantage of the OMOP (Observational Medical Outcomes Partnership) standard, BOAR is able to represent records retrieved from different data models. The experimental results show that BOAR is flexible and robust to representing EHR from several sources, and allows the execution and evaluation of several clustering algorithms. José Maria Clementino, Christian C. Bones, Bruno S. Faiçal, Oscar A. C. Linares, Daniel Mario de Lima, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 8 |
| 2020 | Efficient Segmentation of Cell Nuclei in Histopathological ImagesabstractComputer-aided cell nuclei segmentation in histology images is essential for image analysis. There is a demand for methods that accurately detect cell nuclei in large images. We introduce the FECS method for automatic cell nuclei segmentation in Hematoxylin and Eosin (H&E) stained histology images. Our method accurately segments cell nuclei, even in large images, efficiently. We use bimodal-like histograms to perform image binarization via the fast Otsu algorithm. We introduce a super-pixel based filter for cell nuclei boundary detection. A Gaussian blur filter allows us to identify cell nuclei centers, which are understood as local minima in the individual cell nuclei regions. We have evaluated our method for two publicly available datasets. Out tests have produced average Jaccard index values of 0.963 and 0.914, respectively, supporting a high degree of segmentation accuracy. We have compared our method against a state-of-the-art method; our method produced better results for both datasets. The average processing time of FECS was approximately just one second for images of 1k x 1k pixel resolution and about three minutes for larger images of 15k x 15k pixel resolution. Oscar A. C. Linares, Aurea Soriano-Vargas, Bruno S. Faiçal, Bernd Hamann, Alexandre Todorovic Fabro, Agma J. M. Traina |
CBMS | 6 |
| 2020 | Taking Advantage of Highly-Correlated Attributes in Similarity Queries with Missing Values
Lucas Santiago Rodrigues, Mirela Teixeira Cazzolato, Agma J. M. Traina, Caetano Traina Jr. |
SISAP | 3 |
| 2019 | How to Automatically Identify Regions of Interest in High-Resolution Images of Lung Biopsy for Interstitial Fibrosis DiagnosisabstractAirway-centered Interstitial Fibrosis (ACIF) is a histological pattern of Interstitial lung diseases. Its diagnosis requires a multidisciplinary approach, in which diverse information, such as clinical data, computed tomography data, and lung biopsy data, is analyzed. Biopsy samples are digitized at high-resolution. Of crucial interest are broncho-and bronchiolocentric remodeling with extracellular matrix deposition. To analyze an image, specialists have to explore it at low microscope magnification, select a region of interest and export a smaller specified sub-image to be interpreted at higher magnification. This process is performed several times, requiring hours, becoming a tiresome task. We propose a method to support pathologists to identify specific patterns of ACIF in high-resolution images from lung biopsies. This can be done by a) automatic microscope magnification reduction; b) computing the probability of pixels belonging to high-density regions; c) extracting Local Binary Patterns (LBP) of the high-and low-density regions; and d) visualizing them in color. We have evaluated our method on nine high-resolution lung biopsies. We have tested the LBP features of high-and low-density regions with the kNN algorithm and obtained a classification accuracy of 94.4%, which is the highest one reported in the literature for this type of data. Oscar A. C. Linares, Bruno S. Faiçal, Paulo Barbosa, Bernd Hamann, Alexandre Todorovic Fabro, Agma J. M. Traina |
CBMS | 6 |
| 2019 | 3DBGrowth: Volumetric Vertebrae Segmentation and Reconstruction in Magnetic Resonance ImagingabstractSegmentation of medical images is critical for making several processes of analysis and classification more reliable. With the growing number of people presenting back pain and related problems, the semi-automatic segmentation and 3D reconstruction of vertebral bodies became even more important to support decision making. A 3D reconstruction allows a fast and objective analysis of each vertebrae condition, which may play a major role in surgical planning and evaluation of suitable treatments. In this paper, we propose 3DBGrowth, which develops a 3D reconstruction over the efficient Balanced Growth method for 2D images. We also take advantage of the slope coefficient from the annotation time to reduce the total number of annotated slices, reducing the time spent on manual annotation. We show experimental results on a representative dataset with 17 MRI exams demonstrating that our approach significantly outperforms the competitors and, on average, only 37% of the total slices with vertebral body content must be annotated without losing performance/accuracy. Compared to the state-of-the-art methods, we have achieved a Dice Score gain of over 5% with comparable processing time. Moreover, 3DBGrowth works well with imprecise seed points, which reduces the time spent on manual annotation by the specialist. Jonathan S. Ramos, Mirela Teixeira Cazzolato, Bruno S. Faiçal, Marcello Henrique Nogueira-Barbosa, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 6 |
| 2019 | A Two-Phase Learning Approach for the Segmentation of Dermatological WoundsabstractTissue segmentation in photographs of lower limb chronic ulcers is a non-intrusive approach that supports dermatological analyses. This paper presents 2PLA, a method that combines supervised and unsupervised learning strategies for enhancing the segmentation of dermatological wounds. Given an ulcer photo captured according to a fixed protocol, 2PLA first phase performs a pixelwise classification of points of interest, whereas pre-processing filters are employed for the smoothing of image noise. The cleaned image is further sent to the 2PLA divide-and-conquer second phase. It builds upon SLIC superpixel construction algorithm for dividing the lower limb into regions of interest with well-defined borders, and clusters the superpixels by taking advantage of the similarity-based DBSCAN algorithm. We set up the phases of our method by using a real annotated set of dermatological wounds, and empirical evaluations on representative samples up to 100,000 points showed a compact Multi-Layer Perceptron with Levenberg-Marquardt training algorithm (Cohen-Kappa = .971, Sensitivity = .98, and Specificity = .98) outperformed other classifiers as 2PLA first phase. Additionally, experimental trials on DBSCAN with five distance functions (L1, L2, L∞, Canberra, and BrayCurtis) indicated L1function provided fewer groups in comparison to the competitors, and the number of clusters was an exponential decay to the similarity ratio. Accordingly, we used the elbow criterion for finding the L1-based DBSCAN threshold as 2PLA second phase parameterization. We evaluated the fine-tuned setting of our method over a labeled set of ulcer images, and wounded tissues were segmented within a .05 Mean Absolute Error ratio. These results illustrate the impact of learning parameters on 2PLA as well as the method efficacy for wound segmentation. Wellington S. Silva, Daniel L. Jasbick, Rodrigo Erthal Wilson, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Daniel de Oliveira 0001, Marcos V. N. Bedo |
CBMS | 5 |
| 2019 | UCORM: Indexing Uncorrelated Metric Spaces for Concise Content-Based Retrieval of Medical ImagesabstractThe large amount of medical exams generated by hospitals has a great potential to boost the support for physicians on decision making tasks. This requires efficient and reliable computational systems to retrieve relevant information in real-time. Existing Content-Based Image Retrieval (CBIR) systems rely on Metric Access Methods (MAMs) to speed-up the retrieval task. In this context, images are represented by Feature Extraction Methods (FEMs), according to information such as color or texture. However, MAMs usually index images based on a single FEM. Whenever physicians want to search for similar images using multiple FEMs simultaneously, they need to perform separated queries. In this work, we propose UCORM, an access method capable of indexing images using multiple FEMs by overlapping different metric spaces. UCORM selects the best FEMs to generate a concise yet accurate indexing space. It relies on an interesting use of Pearson correlation, that we named PCMS, to compute the correlation between different FEMs. PCMS allows UCORM to improve the retrieval task by minimizing the overlapping between metric spaces, resulting on fewer intermediary images when performing a query. Experimental analysis shows that UCORM prunes well the data distribution regions with low correlation between FEMs. Also, two medical application scenarios support our claim that UCORM is well-fitted for clinical environments. Guilherme F. Zabot, Mirela Teixeira Cazzolato, Lucas C. Scabora, Bruno S. Faiçal, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 5 |
| 2019 | Efficient Indexing of Multiple Metric Spaces with SpectraabstractThe widespread of social networks and online channels has increased the capture of large amounts of complex data, such as images and videos, which demand efficient and flexible tools to perform information retrieval. Many existing approaches to retrieve complex data follow the "Query by Similarity" paradigm, using Metric Access Methods (MAMs) to index complex data and speed-up information retrieval. In this context, many descriptors represent complex data using representative features such as color, shape, or texture for images. MAMs were initially designed to index features from complex data using only one descriptor, leading users to build several indexes when more than one descriptor is required. Recent approaches that use different representations in a single index structure suffer from a higher number of distance calculations. In this work, we propose the Spectra MAM, which indexes complex data using several features at once. Spectra integrates several metric spaces and answers queries based on one or more descriptors at once. Moreover, Spectra relies on existing correlations among different spaces to choose the best descriptors to obtain a concise yet accurate indexing space. Thus, it reduces the number of distance calculations, speeding up the query execution, and improving the resulting quality. Guilherme F. Zabot, Mirela Teixeira Cazzolato, Lucas C. Scabora, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 4 |
| 2019 | Querying on large and complex databases by content: Challenges on variety and veracity regarding real applications
Agma J. M. Traina, Safia Brinis, Glauco Vitor Pedrosa, Letricia P. S. Avalhais, Caetano Traina Jr. |
Inf. Syst. | 1 |
| 2019 | Hollow-tree: a metric access method for data with missing values
Safia Brinis, Caetano Traina Jr., Agma J. M. Traina |
J. Intell. Inf. Syst. | 3 |
| 2019 | SemIndex+: A semantic indexing scheme for structured, unstructured, and partly structured data
Joe Tekli, Richard Chbeir, Agma J. M. Traina, Caetano Traina Jr. |
Knowl. Based Syst. | 3 |
| 2019 | Employing Domain Indexes to Efficiently Query Medical Data From Multiple RepositoriesabstractContent-based retrieval still remains one of the main problems with respect to controversies and challenges in digital healthcare over big data. To properly address this problem, there is a need for efficient computational techniques, especially in scenarios involving queries across multiple data repositories. In such scenarios, the common computational approach searches the repositories separately and combines the results into one final response, which slows down the process altogether. In order to improve the performance of queries in that kind of scenario, we present the Domain Index, a new category of index structures intended to efficiently query a data domain across multiple repositories, regardless of the repository to which the data belong. To evaluate our method, we carried out experiments involving content-based queries, namely range and k nearest neighbor (kNN) queries, 1) over real-world data from a public data set of mammograms, as well as 2) over synthetic data to perform scalability evaluations. The results show that images from any repository are seamlessly retrieved, sustaining performance gains of up to 53% in range queries and up to 81% in kNN queries. Regarding scalability, our proposal scaled well as we increased 1) the cardinality of data (up to 59% of gain) and 2) the number of queried repositories (up to 71% of gain). Hence, our method enables significant performance improvements, and should be of most importance for medical data repository maintainers and for physicians' IT support. Paulo H. Oliveira, Lucas C. Scabora, Mirela Teixeira Cazzolato, Willian D. Oliveira, Rafael S. Paixão, Agma J. M. Traina, Caetano Traina Jr. |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | ICARUS: Retrieving Skin Ulcer Images through Bag-of-SignaturesabstractThe images collected during medical exams are a strong asset for diagnosing and decision making. One scenario where clinical images are especially useful is the analysis of chronic lesions on the skin (skin ulcers). The visual appearance of these wounds may provide meaningful clues that may help physicians in the diagnosis. In this context, we propose ICARUS, an image retrieval system for dermatological ulcer images based on Bag-of-Visual-Words of color and texture signatures. ICARUS analyzes the image and extracts only the relevant signatures. The results show that ICARUS achieves improvement of up to 7% in image retrieval precision whereas being up to 5 orders of magnitude faster when compared to the state-of-the-art methods. Our results showed that ICARUS is effective and fast, and successfully adds semantic to the image representation. Daniel Y. T. Chino, Lucas C. Scabora, Mirela Teixeira Cazzolato, Ana Elisa Serafim Jorge, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 6 |
| 2018 | RAFIKI: Retrieval-Based Application for Imaging and Knowledge InvestigationabstractMedical exams, such as CT scans and mammograms, are obtained and stored every day in hospitals all over the world, including images, patient data, and medical reports. It is paramount to have tools and systems to improve computer-aided diagnoses based on such huge volumes of stored information. The Content-Based Image Retrieval (CBIR) is a powerful paradigm to help reaching such a goal, providing physicians with intelligent retrieval tools to present him/her with similar or complementary cases, in which visual characteristics improve textual data. Employing comparative inspection on previous cases, the physician can obtain a more comprehensive understanding of the case he/she is working on. Current hospital systems do not carry native CBIR functionalities yet, relying on add-on subsystems, which often do not adhere to the existing relational database infrastructures. In this work, we propose RAFIKI, a software prototype that extends the Relational Database Management System (RDBMS) PostgreSQL, providing native support for CBIR functionalities, modular extensibility, and seamless integration for data science tools, such as Python and R. We show the applicability of our system by evaluating three clinical scenarios, performing queries over a real-world image dataset of lung exams. Our results spot actual potential in promoting informed decision-making from the physician's perspective. Besides, the system exhibited a higher performance when compared to previous systems found in the literature. Moreover, RAFIKI contributes with a model to establish how to put together CBIR concepts and relational data, providing a powerful design for further development of theoretical and practical concepts and tools. Marcos Roberto Nesso Junior, Mirela Teixeira Cazzolato, Lucas C. Scabora, Paulo H. Oliveira, Gabriel Spadon, Jéssica Andressa de Souza, Willian D. Oliveira, Daniel Y. T. Chino, José F. Rodrigues Jr., Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 10 |
| 2018 | Efficient and Reliable Estimation of Cell PositionsabstractSequences of microscopic images feature the dynamics of developing embryos. Automatically tracking the cells from such sequences of images allows understanding the dynamics which a living element demands to know its cells movement, which ideally should take place in real-time. The traditional tracking pipeline starts with image acquisition, data transfer, image segmentation to separate cells from the background, and then the actual tracking step. To speed up this pipeline, we hypothesize that a process capable of predicting the cell motion according to previous observations is useful. The solution must be accurate, fast and lightweight, and be able to iterate between the various components. In this work we propose CM-Predictor, which takes advantage of previous positions of cells to estimate their motion. When estimation takes place, we can omit costly acquisition, transfer and process of images, speeding up the tracking pipeline. The designed solution monitors the error of prediction, adapting the model whenever needed. For validation, we use four different datasets with sequences of images with developing embryos. Then we compare the estimated motion vectors of CM-Predictor with traditional tracking methods. Experimental results show that CM-Predictor is able to accurately estimate the motion vectors. In fact, CM-Predictor maintains the prediction quality of other algorithms and performs faster than them. Mirela Teixeira Cazzolato, Agma J. M. Traina, Klemens Böhm |
CIKM | 2 |
| 2018 | Exploring Diversified Similarity with KundahaabstractExploring large medical image sets by means of traditional similarity query criteria (e.g., neighborhood) can be fruitless if retrieved images are too similar among themselves. This demonstration introduces Kundaha, an exploration tool that assists experts in retrieving and navigating on results from a diversified similarity perspective of user-posed queries. Its implementation includes a wide set of metrics, descriptors, and indexes for enhancing query execution. Users can combine such features with diversified similarity criteria for the organized exploration of result sets and also employ relevance feedback cycles for finding new query-based viewpoints. Lúcio F. D. Santos, Gustavo Blanco, Daniel de Oliveira 0001, Agma J. M. Traina, Caetano Traina Jr., Marcos V. N. Bedo |
CIKM | 4 |
| 2018 | The Merkurion approach for similarity searching optimization in Database Management SystemsabstractModern Database Management Systems (DBMSs) retrieve songs that resemble those in a music dataset, identify plagiarism in a set of documents, or provide past cases to physicians by taking into account the characteristics of a query exam. All such tasks require the comparison of data by similarity, which can be expressed in terms of distance-based queries in metric spaces. Traditional query processing relies mostly on histograms for describing the data distribution space and choosing a data retrieval path that quickly leads to the answer, discarding comparisons of most unwanted data. However, DBMSs still lack adequate support for selectivity estimation of query operators for data types embedded in metric spaces. This article addresses a novel strategy that extends the query optimizer of a DBMS, so that it can also perform both logical and physical query plan optimizations in searches that include similarity predicates. The proposal, named Merkurion, updates the concept of Data Distribution Space and captures data distributions according to the distances between the elements within a dataset. Moreover, it employs concise representations of such distributions, called synopses, for the definition of rules that enable similarity searching optimization. An extensive evaluation of Merkurion in real-world datasets has proven its effectiveness and broad applicability to many data domains. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
Data Knowl. Eng. | 3 |
| 2018 | Full-fledged semantic indexing and querying model designed for seamless integration in legacy RDBMS
Joe Tekli, Richard Chbeir, Agma J. M. Traina, Caetano Traina Jr., Kokou Yétongnon, Carlos Raymundo Ibañez, Marc Al Assad, Christian Kallas |
Data Knowl. Eng. | 3 |
| 2018 | How to speed up outliers removal in image matching
Jonathan S. Ramos, Carolina Yukari Veludo Watanabe, Caetano Traina Jr., Agma J. M. Traina |
Pattern Recognit. Lett. | 4 |
| 2017 | BREATH: Heat Maps Assisting the Detection of Abnormal Lung Regions in CT ScansabstractComputed Tomography (CT) scans are often employed to diagnose lung diseases, as abnormal tissue regions may indicate whether proper treatment is required. However, detecting specific regions containing abnormalities in a CT scan demands time and effort of specialists. Moreover, different parts of a single lung image may present both normal and abnormal characteristics, what makes inaccurate the classification of a single lung as healthy (normal) or not. In this paper we propose the BREATH method, capable of detecting abnormalities in lung tissue regions, highlighting them by means of a heat map visualization. The method starts by segmenting lung tissues using a superpixel-based approach, followed by the training of a statistical model to represent normal tissues and, finally, the generation of a heat map showing abnormal regions that require attention from the physicians. We validated our statistical model using a dataset with 246 lung CT scans, where 40 are healthy and the remaining present varying diseases. Experimental results show that BREATH is accurate for lung segmentation with F-Measure of up to 0.99. The statistical modeling of healthy and abnormal lung regions has shown almost no overlap, and the detection of superpixels containing abnormalities presented precision values higher than 86%, for all values of recall. These values support our claim that the heat map representation of BREATH for the abnormal detection can be used as an intuitive method to assist physicians during the diagnosis. Mirela Teixeira Cazzolato, Lucas C. Scabora, Alceu Ferraz Costa, Marcos Roberto Nesso Junior, Luis Fernando Milano Oliveira, Daniel S. Kaster, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 8 |
| 2017 | Efficiently Indexing Multiple Repositories of Medical Image DatabasesabstractPerforming content-based image retrieval over large repositories of medical images demands efficient computational techniques. The use of such techniques is intended to speed up the work of physicians, who often have to deal with information from multiple data repositories. When dealing with multiple data repositories, the common computational approach is to search each repository separately and merge the multiple results into one final response, which slows down the whole process. This can be improved if we build a mechanism able to search several repositories as if they were a single one, i.e. a mechanism to search the whole domain of medical images. Aiming at this goal, we propose the Domain Index, a new category of index structures aimed at efficiently searching domains of data, regardless of the repository to which they belong. To evaluate our proposal, we carried out experiments over multiple mammography repositories involving k Nearest Neighbor (kNN) and Range queries. The results show that images from any repository are seamlessly retrieved, even sustaining gains in performance of up to 36% in kNN queries and up to 7% in Range queries. The experimental evaluation shows that the Domain Index allows fast retrieval from multiple data repositories for medical systems, allowing a better performance in similarity queries over them. Paulo H. Oliveira, Lucas C. Scabora, Mirela Teixeira Cazzolato, Willian D. Oliveira, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 5 |
| 2017 | VolTime: Unsupervised Anomaly Detection on Users' Online Activity VolumeabstractIs it possible to spot review frauds and spamming on social media and online stores? In this paper we analyze the joint distribution of the inter-arrival times and volume of events such as comments and online reviews and show that it is possible to accurately rank and detect suspicious users such as spammers, bots and fraudsters. We propose VolTime, a generative model that fits well the inter-arrival time distribution (IAT) of real users. Thus, VOLTIME automatically spots and ranks suspicious users. Experiments on several real datasets, ranging from Reddit comments and phone calls to Flipkart product reviews, show that VolTime is able to accurately fit the activity volume and IAT of real data. Additionally, we show that VolTime ranks suspicious users with a precision higher than 90% for a sensitivity of 70%. Daniel Y. T. Chino, Alceu Ferraz Costa, Agma J. M. Traina, Christos Faloutsos |
SDM | 3 |
| 2017 | Semantic Similarity Group By Operators for Metric Data
Natan A. Laverde, Mirela Teixeira Cazzolato, Agma J. M. Traina, Caetano Traina Jr. |
SISAP | 3 |
| 2017 | Retrieving 2D shapes by similarity based on bag of salience points
Glauco Vitor Pedrosa, Agma J. M. Traina, Célia A. Zorzo Barcelos |
Multim. Tools Appl. | 2 |
| 2017 | Modeling Temporal Activity to Detect Anomalous Behavior in Social MediaabstractSocial media has become a popular and important tool for human communication. However, due to this popularity, spam and the distribution of malicious content by computer-controlled users, known as bots, has become a widespread problem. At the same time, when users use social media, they generate valuable data that can be used to understand the patterns of human communication. In this article, we focus on the following important question: Can we identify and use patterns of human communication to decide whether a human or a bot controls a user? The first contribution of this article is showing that the distribution of inter-arrival times (IATs) between postings is characterized by following four patterns: (i) heavy-tails, (ii) periodic-spikes, (iii) correlation between consecutive values, and (iv) bimodallity. As our second contribution, we propose a mathematical model named Act-M (Activity Model). We show that Act-M can accurately fit the distribution of IATs from social media users. Finally, we use Act-M to develop a method that detects if users are bots based only on the timing of their postings. We validate Act-M using data from over 55 million postings from four social media services: Reddit, Twitter, Stack-Overflow, and Hacker-News. Our experiments show that Act-M provides a more accurate fit to the data than existing models for human dynamics. Additionally, when detecting bots, Act-M provided a precision higher than 93% and 77% with a sensitivity of 70% for the Twitter and Reddit datasets, respectively. Alceu Ferraz Costa, Yuto Yamaguchi, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
ACM Trans. Knowl. Discov. Data | 3 |
| 2016 | Encoding Visual Attention Features for Effective Biomedical Images RetrievalabstractThis paper proposes a novel model, called Similarity Based on Visual Attention Features (SimVisual), to enhance the similarity analysis between images by considering features extracted from salient regions mapped by visual attention models. Visual attention models have demonstrated to be very useful for encoding perceptual semantic information of the image content. Thus, aggregating saliency features into the final image representation is a powerful asset to enhance the similarity analysis between images, while increasing the accuracy in retrieval tasks. The goal of SimVisual is to combine different saliency models with traditional image descriptors, aimed at increasing the descriptive power of these descriptors without modifying the original algorithms. We performed some experiments using a large dataset composed of 32 different biomedical images categories, and the results show that SimVisual boosts the retrieval accuracy up to 13% considering simple image descriptors, such as Color Histograms. The experiments on SimVisual shows that it is a valuable approach to increase the efficacy of content-based image retrieval systems, without user interactions. Glauco Vitor Pedrosa, Agma J. M. Traina |
CBMS | 2 |
| 2016 | Vote-and-Comment: Modeling the Coevolution of User Interactions in Social Voting Web SitesabstractIn social voting Web sites, how do the user actions - up-votes, down-votes and comments - evolve over time? Are there relationships between votes and comments? What is normal and what is suspicious? These are the questions we focus on. We analyzed over 20,000 submissions corresponding to more than 100 million user interactions from three social voting Web sites: Reddit, Imgur and Digg. Our first contribution is two discoveries: (i) the number of comments grows as a power-law on the number of votes and (ii) the time between a submission creation and a user's reaction obeys a log-logistic distribution. Based on these patterns, we propose VnC (Vote-and-Comment), a parsimonious but accurate and scalable model that models the coevolution of user activities. In our experiments on real data, VnC outperformed state-of-the-art baselines on accuracy. Additionally, we illustrate VnC usefulness for forecasting and outlier detection. Alceu Ferraz Costa, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
ICDM | 2 |
| 2016 | Fire Detection on Unconstrained Videos Using Color-Aware Spatial Modeling and Motion FlowabstractThe semantic segmentation of events on emergency contexts involves the identification of previously defined events of interest. In this work, the focused semantic event is the presence of fire in videos. The literature presents several methods for automatic video fire detection, but these methods were built under assumptions, such as stationary cameras and controlled lightening conditions that are often in contrast to the videos acquired by hand-held devices. To fulfill this gap, we propose a fire detection method, called SPATFIRE. Our method innovates on three aspects: (1) it relies on a specifically tailored color model named Fire-like Pixel Detector able to improve the accuracy of fire detection, (2) it employs a new technique for motion compensation, diminishing the problems observed in videos captured with non-stationary cameras, and, (3) it defines a segmentation method able to identify, not only the presence of fire in a video, but also the segments in the video where fire occurs. We experimented our proposal on two video datasets with different characteristics and summarize the results to demonstrate the superior efficacy, in terms of true positives and negatives, as compared to state-of-the-art methods. Letricia P. S. Avalhais, José F. Rodrigues Jr., Agma J. M. Traina |
ICTAI | 3 |
| 2016 | A Label-Scaled Similarity Measure for Content-Based Image RetrievalabstractContent-Based Image Retrieval (CBIR) has proven to be a suitable complement to traditional text-based searching. CBIR applications rely on two main steps, namely the representation of the images, and the similarity measuring between two represented images. Although modern segmentation and learning algorithms enable the accurate representation of local and global features within an image, how to properly compare the segmented objects is still an open issue. In this study, we propose a new comparison method called Counting-Labels Similarity Measure (CL-Measure). Our approach calculates the similarity between two images by comparing the labeled regions within these images and by balancing the influence of each label according to its predominance in both non-metric and metric fashion. The experiments on a real dataset of dermatological ulcers show that CL-Measure achieves a higher Precision for all values of Recall compared to its competitors in retrieval tasks. Gustavo Blanco, Marcos V. N. Bedo, Mirela Teixeira Cazzolato, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina |
ISM | 8 |
| 2016 | When Similarity is Not Enough, Ask for Diversity: Grouping Elements Based on InfluenceabstractCrowdsourcing images have been increasingly employed for mapping emergency scenarios, which helps rescue forces in choosing contingency plans. In this scenario, similarity searching can be used to retrieve related images from past situations. However, the retrieved images often are similar among themselves and, therefore, add little to none new information to the rescue decision-making process. In this paper, we take advantage of diversity queries to increase the variety of the representative elements about an incident, whereas the remaining and related data are grouped according to the set of representatives. Thus, our approach enables content retrieval, grouping and an easier exploration of the result set. Experiments performed on real datasets shows that our proposal outperforms the existing methods regarding both quality and performance, being at least three orders of magnitude faster. Lúcio F. D. Santos, Luiz Olmes Carvalho, Marcos V. N. Bedo, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 4 |
| 2016 | Preface
Agma J. M. Traina, Caetano Traina Jr. |
Inf. Syst. | 1 |
| 2015 | Vertebral Body Segmentation of Spine MR Images Using SuperpixelsabstractThis paper presents a segmentation approach guided by the user for extracting the vertebral bodies of spine from MRI. The proposed approach, called VBSeg, takes advantage of super pixels to reduce the image complexity and then making easy the detection of each vertebral body contour. Super pixels adapt themselves to the image structures, once their formation law follows the homogeneity of the image regions. However, for some diseases or abnormalities, the boundary of each super pixel does not fit well in the vertebra contour. To avoid this drawback, we propose to use the Otsu's method as a possegmentation step to divide the super pixels into smaller ones. The final segmentation is obtained through a region growing approach using points manually selected by the specialist. It can produce masks of the five lumbar vertebrae with an average precision of 80% and recall of 87%, when compared to the manual segmentation of a trained specialist. These values show that the VBSeg is a valuable asset to assist the medical specialist in the task of vertebral bodies' segmentation, with much less effort and time demand. Paulo Duarte Barbieri, Glauco Vitor Pedrosa, Agma J. M. Traina, Marcello Henrique Nogueira-Barbosa |
CBMS | 3 |
| 2015 | Color and Texture Influence on Computer-Aided Diagnosis of Dermatological UlcersabstractThis study presents an analysis of classification techniques for Computer-Aided Diagnosis (CAD) regarding ulcerated lesions. We focus on determining influence of both color and texture in the automated image classification and its implication. To do so, we assayed a dataset of dermatological ulcers containing five variations in terms of tissue composition of lesion skin: granulation (red), fibrin (yellow), callous (white), necrotic (black), and a mix of the previous variations (mixed). Every image was previously labelled by experts regarding this red-yellow-black-white-mixed model. We employed specially designed color and texture extractors to represent the dataset images, namely: Color Layout, Color Structure, Scalable Color, Edge Histogram, Haralick, and Texture-Spectrum. The first three are color feature extractors and the last three are texture extractors. Following, we employed the Symmetrica Uncert Attribute Eval method to determine the features suitable for image classification. We tested a set of classifiers that follows distinct paradigms over the selected features, achieving an accuracy ratio of up to 77% in terms of images correctly classified, with the area under the receiver operating characteristic (ROC) curve up to 0.84. The classification performance and the selected features enabled us to determine that texture features were more predominant than color in the entire classification process. Marcos V. N. Bedo, Lúcio F. D. Santos, Willian D. Oliveira, Gustavo Blanco, Agma J. M. Traina, Marco Antonio Frade, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr. |
CBMS | 5 |
| 2015 | Speeding up the combination of multiple descriptors for different boundary conditionsabstractContent-based complex data retrieval is becoming increasingly common in many types of applications. The content of these data is represented by intrinsic characteristics, extracted from them which together with a distance function allows similarity queries. Aimed at reducing the "semantic gap", characterized by the disagreement between the computational representation of the extracted low-level features and how these data are interpreted by the human perception, the use of multiple descriptors has been the subject of several studies. This paper proposes a new method to carry out the combination of multiple descriptors for different boundary conditions in which the balancing is carried out in pairs, starting by the best candidate descriptor. In the experiments, the proposed method achieved computational cost up to 3650 times smaller than the exhaustive search for the best linear combination of descriptors, keeping almost the same average precision, with variations lower than 0.9%. Rodrigo Fernandes Barroso, Marcelo Ponciano-Silva, Agma J. M. Traina, Renato Bueno |
CLEI | 3 |
| 2015 | Self Similarity Wide-Joins for Near-Duplicate Image DetectionabstractNear-duplicate image detection plays an important role in several real applications. Such task is usually achieved by applying a clustering algorithm followed by refinement steps, which is a computationally expensive process. In this paper we introduce a framework based on a novel similarity join operator, which is able both to replace and speed up the clustering step, whereas also releasing the need of further refinement processes. It is based on absolute and relative similarity ratios, ensuring that top ranked image pairs are in the final result. Experiments performed on real datasets shows that our proposal is up to three orders of magnitude faster than the best techniques in the literature, always returning a high-quality result set. Luiz Olmes Carvalho, Lúcio F. D. Santos, Willian D. Oliveira, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 4 |
| 2015 | Combining Diversity Queries and Visual Mining to Improve Content-Based Image Retrieval Systems: The DiVI MethodabstractThis paper proposes a new approach to improve similarity queries with diversity, the Diversity and Visually-Interactive method (DiVI), which employs Visual Data Mining techniques in Content-Based Image Retrieval (CBIR) systems. DiVI empowers the user to understand how the measures of similarity and diversity affect their queries, as well as increases the relevance of CBIR results according to the user judgment. An overview of the image distribution in the database is shown to the user through multidimensional projection. The user interacts with the visual representation changing the projected space or the query parameters, according to his/her needs and previous knowledge. DiVI takes advantage of the users' activity to transparently reduce the semantic gap faced by CBIR systems. Empirical evaluation show that DiVI increases the precision for querying by content and also increases the applicability and acceptance of similarity with diversity in CBIR systems. Lúcio F. D. Santos, Rafael L. Dias, Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 4 |
| 2015 | RSC: Mining and Modeling Temporal Activity in Social MediaabstractCan we identify patterns of temporal activities caused by human communications in social media? Is it possible to model these patterns and tell if a user is a human or a bot based only on the timing of their postings? Social media services allow users to make postings, generating large datasets of human activity time-stamps. In this paper we analyze time-stamp data from social media services and find that the distribution of postings inter-arrival times (IAT) is characterized by four patterns: (i) positive correlation between consecutive IATs, (ii) heavy tails, (iii) periodic spikes and (iv) bimodal distribution. Based on our findings, we propose Rest-Sleep-and-Comment (RSC), a generative model that is able to match all four discovered patterns. We demonstrate the utility of RSC by showing that it can accurately fit real time-stamp data from Reddit and Twitter. We also show that RSC can be used to spot outliers and detect users with non-human behavior, such as bots. We validate RSC using real data consisting of over 35 million postings from Twitter and Reddit. RSC consistently provides a better fit to real data and clearly outperform existing models for human dynamics. RSC was also able to detect bots with a precision higher than 94%. Alceu Ferraz Costa, Yuto Yamaguchi, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
KDD | 3 |
| 2015 | Similarity Joins and Beyond: An Extended Set of Binary Operators with Order
Luiz Olmes Carvalho, Lúcio F. D. Santos, Willian D. Oliveira, Agma J. M. Traina, Caetano Traina Jr. |
SISAP | 4 |
| 2015 | Improving Metric Access Methods with Bucket Files
Ives Rene Venturini Pola, Agma J. M. Traina, Caetano Traina Jr., Daniel S. Kaster |
SISAP | 2 |
| 2015 | Diversity in Similarity Joins
Lúcio F. D. Santos, Luiz Olmes Carvalho, Willian D. Oliveira, Agma J. M. Traina, Caetano Traina Jr. |
SISAP | 4 |
| 2015 | Compact distance histogram: a novel structure to boost k-nearest neighbor queriesabstractThe k-Nearest Neighbor query (k-NNq) is one of the most useful similarity queries. Elaborated k-NNq algorithms depend on an initial radius to prune regions of the search space that cannot contribute to the answer. Therefore, estimating a suitable starting radius is of major importance to accelerate k-NNq execution. This paper presents a new technique to estimate a tight initial radius. Our approach, named CDH-kNN, relies on Compact Distance Histograms (CDHs), which are pivot-based histograms defined as piecewise linear functions. Such structures approximate the distance distribution and are compressed according to a given constraint, which can be a desired number of buckets and/or a maximum allowed error. The covering radius of a k-NNq is estimated based on the relationship between the query element and the CDHs' joint frequencies. The paper presents a complete specification of CDH-kNN, including CDH's construction and radii estimation. Extensive experiments on both real and synthetic datasets highlighted the efficiency of our approach, showing that it was up to 72% faster than existing algorithms, outperforming every competitor in all the setups evaluated. In fact, the experiments showed that our proposal was just 20% slower than the theoretical lower bound. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 3 |
| 2015 | Similarity sets: A new concept of sets to seamlessly handle similarity in database management systems
Ives Rene Venturini Pola, Robson L. F. Cordeiro, Caetano Traina Jr., Agma J. M. Traina |
Inf. Syst. | 4 |
| 2015 | Approximate XML structure validation based on document-grammar tree similarity
Joe Tekli, Richard Chbeir, Agma J. M. Traina, Caetano Traina Jr., Renato Fileto |
Inf. Sci. | 3 |
| 2014 | SemIndex: Semantic-Aware Inverted Index
Richard Chbeir, Joe Tekli, Kokou Yétongnon, Carlos Raymundo Ibañez, Agma J. M. Traina, Caetano Traina Jr., Marc Al Assad |
ADBIS | 6 |
| 2014 | MedInject: A General-Purpose Information Retrieval Framework Applied in a Medical ContextabstractThe continuous improvement of medical software and instrumentation have contributed to generate large amounts of medical image data. Thus, plenty of Content-Based Image Retrieval systems have emerged in order to index and retrieve images according to similarity criteria. Some of those systems are applied in very specific domains, such as mammography, lung or spine exams. Others, however, are general-purpose applications that can be adopted in a medical environment. In such context, we realized those specific systems could benefit from the facilities brought by generic frameworks and propose our solution. This article presents a novel information retrieval core framework that performs both indexing and similarity search operations over medical image data sets. The framework follows a modular architecture based on Design Patterns and can be easily extended, allowing to other system developers to take advantages of its functions by using the provided interfaces. We performed extensive experiments evaluating several of its properties and target abstractions using medical real data, and show that it allows the implementation to achieve proper similarity retrieval and significant performance improvements in relation to the existing alternatives. Luiz Olmes Carvalho, Enzo Seraphim, Thatyana F. P. Seraphim, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 4 |
| 2014 | Using Sub-dictionaries for Image Representation Based on the Bag-of-Visual-Words ApproachabstractBag-of-Visual-Words (BoVW) is a well known approach to represent images for visual recognition and retrieval tasks. This approach represents an image as a histogram of visual words and the dissimilarity between two images is measured by comparing those histograms. When performing comparisons involving a specific type of images, some visual words can be more informative and discriminative than others. To take advantage of this fact, assigning appropriate weights can improve the performance of image retrieval. In this paper, we developed a novel modeling approach based on sub dictionaries. We extracted a sub-dictionary as a subset of visual words that best represents a specific image class. To measure the dissimilarity distance between images, we take into account the distance of the histogram obtained using the visual dictionary and the distances of the sub histograms obtained by each sub-dictionary. The proposed approach was evaluated by classifying a standard biomedical image dataset into categories defined by image modality and body part and also natural image scenes. The experimental results demonstrate the gain obtained of the proposed weighting approach when compared to the traditional weighting approach based on TF-IDF (Term Frequency-Inverse Document Frequency). Our proposed approach has shown promising results to boost the classification accuracy as well as the retrieval precision. Moreover, it does that without increasing the feature vector dimensionality. Glauco Vitor Pedrosa, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2014 | Being Similar is Not Enough: How to Bridge Usability Gap through Diversity in Medical ImagesabstractIn this paper we present a technique developed to bridge the usability gap in Content-Based Medical Image Retrieval (CBMIR) systems exploring both similarity and diversity. Usability gaps are related to how easy to use a software tool from the radiologist's perspective is. Although much have been done to better express similarity queries, the use of CBMIR over massive databases may have drawbacks that impact its usability. We claim that much of the problems derives from the fact that many images returned are closer to each other than to the query element (near-duplicates). To target this nuisance, we propose to boost similarity queries with diversity, using a technique to hierarchically cluster near-duplicates. We tailored a domain-independent and parameter-free method by controlling the maximum area reached in the search space. This novel approach to improve CBMIR systems take advantage of diversity expectations. The proposed approach BridGE (Better result with influence diversification to Group Elements) aims at adding new relevant information to the analysts, reducing the need of further query refinement or relevance feedback cycles. The results are displayed to the specialist as a traditional CBMIR result whereas the radiologists are able to expand the clusters and navigate through them. The results support our claim that a CBMIR system empowered with diversity is able to bridge the usability gap, grouping near-duplicates and being at least 2 orders of magnitude faster than its mainly competitors. Lúcio F. D. Santos, Marcos V. N. Bedo, Marcelo Ponciano-Silva, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 4 |
| 2014 | Integrating visual words as bunch of n-grams for effective biomedical image classificationabstractThe Bag-of-Visual-Words (BoVW) has been frequently used in the classification of image data. However, this modeling approach does not take into consideration the spatial relationships of these words, which is important for similarity measurement between images. We have developed a novel technique to incorporate spatial information of visual words based on the n-grams representation. The method encodes regional layout with a 2-gram representation in the local keypoint neighborhood. The region is divided in two zones to capture the relative orientations of pair-wise visual words. In turn, each image is described by an accumulated vector of 2-grams. Then, we compute the Shannon entropy over a random “bunch” of 2-grams to reduce the dimensionality of the feature vector. We discovered that this reduction technique creates a more discriminative feature vector as well as presents a considerable dimensionality reduction of up to 99%. The final representation is a compact and efficient local image descriptor that encodes frequency and arrangement of visual words. The proposed approach was tested by classifying a standard biomedical image dataset into categories defined by image modality and body part. The experimental results demonstrate the importance of contextual relations of visual words. Our proposed approach improved the classification accuracy compared to the traditional BoVW by 6.03%. Glauco Vitor Pedrosa, Sameer K. Antani, Dina Demner-Fushman, L. Rodney Long, Agma J. M. Traina |
WACV | 6 |
| 2014 | The NOBH-tree: Improving in-memory metric access methods by using metric hyperplanes with non-overlapping nodes
Ives Rene Venturini Pola, Caetano Traina Jr., Agma J. M. Traina |
Data Knowl. Eng. | 3 |
| 2014 | QuMinS: Fast and scalable querying, mining and summarizing multi-modal databases
Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Luciana A. S. Romani, Priscila P. Coltri, Tamires T. Souza, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
Inf. Sci. | 10 |
| 2013 | Does a CBIR system really impact decisions of physicians in a clinical environment?abstractContent-based image retrieval systems are employed in several areas. One of the most prominent area is the medical field, due to the huge volume of digital images daily generated in healthcare institutions employed for decision making. There are several works applying CBIR techniques over medical images. However, the great majority of them do not verify whether the systems are actually considered by the specialists as a pontential aid in a real environment. In order to fill this research void in the literature, this work explores user experiments in a CBIR system involving resident physicians and radiologists. To do so, we developed a CBIR system according to requirements provided by the specialists and employed a methodology to analyze the effectiveness of the system for supporting them in clinical routine. The methodology aims at evaluating the system's impact in the user's decision, inquiring the specialists about the image classification and their degree of certainty in different situations using the system. By analyzing the obtained results we can argue that the proposed methodology joined with our medical CBIR system presented a high acceptance and viability rate regarding the radiologists interests in the clinical practice domain, providing a novel approach to analyze CBIR systems under realistic conditions. Marcelo Ponciano-Silva, Juliana P. Souza, Pedro Henrique Bugatti, Marcos V. N. Bedo, Daniel S. Kaster, Rosana T. V. Braga, Angela D. Bellucci, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 10 |
| 2013 | Using Boundary Conditions for Combining Multiple Descriptors in Similarity Based Queries
Rodrigo Fernandes Barroso, Marcelo Ponciano-Silva, Agma J. M. Traina, Renato Bueno |
CIARP (1) | 3 |
| 2013 | A Differential Method for Representing Spinal MRI for Perceptual-CBIR
Marcelo Ponciano-Silva, Pedro Henrique Bugatti, Rafael Menezes-Reis, Paulo Mazzoncini de Azevedo Marques, Marcello Henrique Nogueira-Barbosa, Caetano Traina Jr., Agma J. M. Traina |
CIARP (1) | 7 |
| 2013 | Efficient Execution of Conjunctive Complex Queries on Big Multimedia DatabasesabstractThis paper proposes an approach to efficientlyexecute conjunctive queries on big complex data together withtheir related conventional data. The basic idea is to horizontallyfragment the database according to criteria frequently usedin query predicates. The collection of fragments is indexed toefficiently find the fragment(s) whose contents satisfy some querypredicate(s). The contents of each fragment are then indexed aswell, to support efficient filtering of the fragment data according to other query predicate s) conjunctively connected to the former. This strategy has been applied to a collection of more than 106 million images together with their related conventional data. Experimental results show considerable performance gain of the proposed approach for queries with conventional and similaritybasedpredicates, compared to the use of a unique metric index for the entire database contents. Karina Fasolin, Renato Fileto, Marcelo Krüger, Daniel S. Kaster, Mônica Ribeiro Porto Ferreira, Robson L. F. Cordeiro, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 7 |
| 2013 | Graph-Based Relational Data VisualizationabstractRelational databases are rigid-structured data sources characterized by complex relationships among a set of relations (tables). Making sense of such relationships is a challenging problem because users must consider multiple relations, understand their ensemble of integrity constraints, interpret dozens of attributes, and draw complex SQL queries for each desired data exploration. In this scenario, we introduce a twofold methodology, we use a hierarchical graph representation to efficiently model the database relationships and, on top of it, we designed a visualization technique for rapidly relational exploration. Our results demonstrate that the exploration of databases is deeply simplified as the user is able to visually browse the data with little or no knowledge about its structure, dismissing the need for complex SQL queries. We believe our findings will bring a novel paradigm in what concerns relational data comprehension. Daniel Mario de Lima, José F. Rodrigues Jr., Agma J. M. Traina |
IV | 3 |
| 2013 | A New Concept of Sets to Handle Similarity in Databases: The SimSets
Ives Rene Venturini Pola, Robson L. F. Cordeiro, Caetano Traina Jr., Agma J. M. Traina |
SISAP | 4 |
| 2013 | Parameter-free and domain-independent similarity search with diversityabstractNew operators to execute similarity-based queries over multimedia data stored in Database Management Systems are increasingly demanded. However, searching in very large datasets, the basic operators often return elements too much similar both to the query center and to themselves, reducing the answer's utility. In this paper, we tackle the problem of providing diversity to similarity query results, and define techniques to assure that each element in the result set is different enough from the others. Existing techniques compel the user to define either a parameter to trade among similarity and diversity or a minimum similarity between result elements. Distinctly, our approach provides similarity queries with diversification using the influence concept, which automatically estimates the inherent diversity between the result set elements requiring no user-defined parameters. Furthermore, our technique can be applied over any data represented in a metric space, so it is both parameter and application-domain independent. The "Better Results with Influence Diversification" (BRID) technique is the basis to the k-Diverse Nearest Neighbor (BRIDk) and to the Range Diverse (BRIDr) algorithms, which execute k-nearest neighbor and range queries with diversification, showing that the technique can be applied to diversify any type of similarity queries. We also define a way to measure the diversification degree in a result set. Through a detailed experimental evaluation using our approach, we show that BRID outperforms the existing methods regarding both query diversification quality and execution times, being at least two orders of magnitude faster than the best existing approaches. Lúcio F. D. Santos, Willian D. Oliveira, Mônica Ribeiro Porto Ferreira, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 4 |
| 2013 | A New Time Series Mining Approach Applied to Multitemporal Remote Sensing ImageryabstractIn this paper, we present a novel unsupervised algorithm, called CLimate and rEmote sensing Association patteRns Miner, for mining association patterns on heterogeneous time series from climate and remote sensing data integrated in a remote sensing information system developed to improve the monitoring of sugar cane fields. The system, called RemoteAgri, consists of a large database of climate data and low-resolution remote sensing images, an image preprocessing module, a time series extraction module, and time series mining methods. The preprocessing module was projected to perform accurate geometric correction, what is a requirement particularly for land and agriculture applications of satellite images. The time series extraction is accomplished through a graphical interface that allows easy interaction and high flexibility to users. The time series mining method transforms series to symbolic representation in order to identify patterns in a multitemporal satellite images and associate them with patterns in other series within a temporal sliding window. The validation process was achieved with agroclimatic data and NOAA-AVHRR images of sugar cane fields. Results show a correlation between agroclimatic time series and vegetation index images. Rules generated by our new algorithm show the association patterns in different periods of time in each time series, pointing to a time delay between the occurrences of patterns in the series analyzed, corroborating what specialists usually forecast without having the burden of dealing with many data charts. Luciana A. S. Romani, Ana Maria Heuminski de Ávila, Daniel Y. T. Chino, Jurandir Zullo Jr., Richard Chbeir, Caetano Traina Jr., Agma J. M. Traina |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2013 | Halite: Fast and Scalable Multiresolution Local-Correlation ClusteringabstractThis paper proposes Halite, a novel, fast, and scalable clustering method that looks for clusters in subspaces of multidimensional data. Existing methods are typically superlinear in space or execution time. Halite's strengths are that it is fast and scalable, while still giving highly accurate results. Specifically the main contributions of Halite are: 1) Scalability: it is linear or quasi linear in time and space regarding the data size and dimensionality, and the dimensionality of the clusters' subspaces; 2) Usability: it is deterministic, robust to noise, doesn't take the number of clusters as an input parameter, and detects clusters in subspaces generated by original axes or by their linear combinations, including space rotation; 3) Effectiveness: it is accurate, providing results with equal or better quality compared to top related works; and 4) Generality: it includes a soft clustering approach. Experiments on synthetic data ranging from five to 30 axes and up to 1 \rm million points were performed. Halite was in average at least 12 times faster than seven representative works, and always presented highly accurate results. On real data, Halite was at least 11 times faster than others, increasing their accuracy in up to 35 percent. Finally, we report experiments in a real scenario where soft clustering is desirable. Robson L. F. Cordeiro, Agma J. M. Traina, Christos Faloutsos, Caetano Traina Jr. |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Large Graph Analysis in the GMine SystemabstractCurrent applications have produced graphs on the order of hundreds of thousands of nodes and millions of edges. To take advantage of such graphs, one must be able to find patterns, outliers, and communities. These tasks are better performed in an interactive environment, where human expertise can guide the process. For large graphs, though, there are some challenges: the excessive processing requirements are prohibitive, and drawing hundred-thousand nodes results in cluttered images hard to comprehend. To cope with these problems, we propose an innovative framework suited for any kind of tree-like graph visual design. GMine integrates 1) a representation for graphs organized as hierarchies of partitions-the concepts of SuperGraph and Graph-Tree; and 2) a graph summarization methodology-CEPS. Our graph representation deals with the problem of tracing the connection aspects of a graph hierarchy with sub linear complexity, allowing one to grasp the neighborhood of a single node or of a group of nodes in a single click. As a proof of concept, the visual environment of GMine is instantiated as a system in which large graphs can be investigated globally and locally. José F. Rodrigues Jr., Hanghang Tong, Jia-Yu Pan, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | k-Gabor: A new feature extraction method for medical images providing internal analysisabstractThis paper proposes the k-Gabor method, a new image feature extractor that captures texture information from medical image regions without a costly segmentation usually associated to texture extractors. It employs Gabor filters, thus, the k-Gabor method can quantify texture information from specific regions, tissues and internal structures of the images providing a succint representation for a richer image analysis. The feature vectors generated describe the images more precisely than other methods from the literature, as shown in the experiments. Besides providing meaningful information from the images, the cost to obtain it is very small, since the total time to extract the k-Gabor features was always only fractions of seconds. Gabriel Humpire-Mamani, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2012 | SART: A New Association Rule Method for Mining Sequential Patterns in Time Series of Climate Data
Marcos Daniel Cano, Marilde Terezinha Prado Santos, Ana Maria Heuminski de Ávila, Luciana A. S. Romani, Agma J. M. Traina, Marcela X. Ribeiro |
ICCSA (3) | 5 |
| 2012 | A Statistical Associative Classifier with Automatic Estimation of Parameters on Computer Aided DiagnosisabstractIn this paper, we proposed a classifier based on statistical association rules that avoids the discretization step and automatically estimates the input thresholds. The algorithm automatically selects the most significant features to produce rules. These rules are simple, including the selected features, a single interval in the antecedent of the rule and a label class in the consequent, and getting at most twice the number of rules features. To evaluate our method, we compare it with traditional classifiers as C4.5 and Adaboost, in the task of classifying benign or malign masses of mammograms, usingtwo different real datasets. The proposed method achieve the best results regarding accuracy, sensitivity and sensibility. Carolina Yukari Veludo Watanabe, Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr. |
ICMLA (1) | 3 |
| 2011 | The Similarity Cloud Model: A novel and efficient hippocampus segmentation techniqueabstractThis work presents a new segmentation model called Similarity Cloud Model (SCM) based on hippocampus feature extraction. The segmentation process is divided in two main operations: localization by similarity and cloud adjustment. The first process uses the cloud to localize the most probable position of the hippocampus in a target volume. Segmentation is completed by a reformulation of the cloud to correct the final labeling, based on a new computation of arc-weights. This method has been tested in an entire dataset of 235 MRI combining healthy and epileptic patients. Results indicate superior quality segmentation in comparison with similar graph and bayesian-based models. Fredy E. C. Atho, Agma J. M. Traina, Caetano Traina Jr., Paula Diniz |
CBMS | 2 |
| 2011 | Improving content-based retrieval of medical images through dynamic distance on relevance feedbackabstractContent-based image retrieval approaches rely on automatic features extracted from images to perform similarity queries. The major drawback is that such features often do not satisfactorily represent what the users understand and expect from them, e.g. when searching for similar images. In order to deal with the gap between the user semantic interpretation of the images and what the system can automatically provide, relevance feedback techniques have been employed. However, it has been used without prior analysis about the distance function that best suits the user intention in each relevance feedback cycle and leading to the increase of such gap. Hence, in the present paper we employ user profiling in conjunction to content-based image retrieval and relevance feedback techniques to exploit the user intentions and to reach the best configuration according to the user intention in each relevance feedback cycle. To do so, we introduce a novel approach and a mediator architecture to enhance this process through user feedback and profiling, allowing to dynamically modify the distance function in each feedback cycle choosing the best one for each cycle according to the user expectation. Experiments have shown that the proposed method outperformed the traditional static distance approach, improving in up to 74% the precision of similarity search of medical images. Pedro Henrique Bugatti, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2011 | Using Visual Analysis to Weight Multiple Signatures to Discriminate Complex DataabstractComplex data is usually represented through signatures, which are sets of features describing the data content. Several kinds of complex data allow extracting different signatures from an object, representing complementary data characteristics. However, there is no ground truth of how balancing these signatures to reach an ideal similarity distribution. It depends on the analyst intent, that is, according to the job he/she is performing, a few signatures should have more impact in the data distribution than others. This work presents a new technique, called Visual Signature Weighting (ViSW), which allows interactively analyzing the impact of each signature in the similarity of complex data represented through multiple signatures. Our method provides means to explore the tradeoff of prioritizing signatures over the others, by dynamically changing their weight relation. We also present case studies showing that the technique is useful for global dataset analysis as well as for inspecting subspaces of interest. Renato Bueno, Daniel S. Kaster, Humberto Luiz Razente, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr. |
IV | 5 |
| 2011 | Clustering very large multi-dimensional datasets with MapReduceabstractGiven a very large moderate-to-high dimensionality dataset, how could one cluster its points? For datasets that don't fit even on a single disk, parallelism is a first class option. In this paper we explore MapReduce for clustering this kind of data. The main questions are (a) how to minimize the I/O cost, taking into account the already existing data partition (e.g., on disks), and (b) how to minimize the network cost among processing nodes. Either of them may be a bottleneck. Thus, we propose the Best of both Worlds -- BoW method, that automatically spots the bottleneck and chooses a good strategy. Our main contributions are: (1) We propose BoW and carefully derive its cost functions, which dynamically choose the best strategy; (2) We show that BoW has numerous desirable features: it can work with most serial clustering methods as a plugged-in clustering subroutine, it balances the cost for disk accesses and network accesses, achieving a very good tradeoff between the two, it uses no user-defined parameters (thanks to our reasonable defaults), it matches the clustering quality of the serial algorithm, and it has near-linear scale-up; and finally, (3) We report experiments on real and synthetic data with billions of points, using up to 1,024 cores in parallel. To the best of our knowledge, our Yahoo! web is the largest real dataset ever reported in the database subspace clustering literature. Spanning 0.2 TB of multi-dimensional data, it took only 8 minutes to be clustered, using 128 cores. Robson L. F. Cordeiro, Caetano Traina Jr., Agma J. M. Traina, Julio López 0002, U Kang, Christos Faloutsos |
KDD | 3 |
| 2011 | Improving the ranking quality of medical image retrieval using a genetic feature selection method
Sérgio Francisco da Silva, Marcela X. Ribeiro, João Batista Neto, Caetano Traina Jr., Agma J. M. Traina |
Decis. Support Syst. | 5 |
| 2011 | Slicing the metric space to provide quick indexing of complex data in the main memory
Caio César Mori Carélo, Ives Rene Venturini Pola, Ricardo Rodrigues Ciferri, Agma J. M. Traina, Caetano Traina Jr., Cristina Dutra de Aguiar Ciferri |
Inf. Syst. | 4 |
| 2010 | Improving medical image retrieval through multi-descriptor similarity functions and association rulesabstractContent-based image retrieval (CBIR) systems still face the problem of low precision of system results. To improve the precision of such systems, many image visual extractors have been developed and employed to represent the images. However, the usage of a large number of extractors and consequently, a large number of features, leads to the “dimensionality curse”, where the retrieval performance and the query accuracy diminish. In this paper, we propose a new method, called Statistical Fractal-scaled Product Metric (SFPM), to maximize the accuracy of CBIR systems and speedup similarity queries. The SFPM method combines association rule mining and the Fractal-scaled Product Metric (FPM) [4], to determine a reduced set of features and appropriate scale factors in multi-descriptor image similarity assessment. The FPM is an unsupervised method to determine a scale factor among features in multi-descriptor image similarity assessment based on the Fractal Theory. Experiments have shown that SFPM reduced the feature vector size in up to 65% and improved in up to 27% the query precision when comparing with the use of the FPM technique. The results show that the proposed method SFPM is effective in determining a reduced set of features and a near-optimal set of scale factors for the descriptors involved, and it is well-suited to improve the quality of content-based query in CBIR systems. Renato Bueno, Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 3 |
| 2010 | Integrating user preference to similarity queries over medical images datasetsabstractLarge amounts of images from medical exams are being stored in databases, so developing retrieval techniques is an important research problem. Retrieval based on the image visual content is usually better than using textual descriptions, as they seldom gives every nuances that the user may be interested in. Content-based image retrieval employs the similarity among images for retrieval. However, similarity is evaluated using numeric methods, and they often orders the images by similarity in a way rather distinct from the user's intention. In this paper, we propose a technique to allow expressing the user's preference over attributes associated to the images, so similarity queries can be refined by preference rules. Experiments performed over a dataset with computed tomography lung images shows that correctly expressing the user's preferences, the similarity query precision can increase from an average of 60% up to close to 100%, when enough interesting images exists in the database. Mônica Ribeiro Porto Ferreira, Marcelo Ponciano-Silva, Agma J. M. Traina, Caetano Traina Jr., Sandra de Amo, Fabíola S. F. Pereira, Richard Chbeir |
CBMS | 3 |
| 2010 | Silhouette-based feature selection for classification of medical imagesabstractClassification is an important task for computer-aided diagnosis systems (CADs). However, many classifiers may not perform well, presenting poor generalization and high computational cost, especially when dealing with high-dimensional datasets. Thus, feature selection can greatly mitigate these problems. In this paper, we propose two filter-based feature selection algorithms that calculate the simplified silhouette statistic as evaluation function: the silhouette-based greedy search (SiGS) and the silhouette-based genetic algorithm search (SiGAS). Silhouette statistic is used to guide the search for features that provide better class separability. Experiments performed on three datasets have shown that the SiGAS algorithm overcomes traditional filter algorithms, such as CFS, FCBF and reliefF. It also outperforms a similar algorithm, kNNGAS, based on genetic algorithm that minimizes the classification error of k-nearest neighbors. Additionally, results have shown that SiGAS produces better accuracy than SiGS. Sérgio Francisco da Silva, Bruno Brandoli Machado, Danilo Medeiros Eler, João Batista Neto, Agma J. M. Traina |
CBMS | 5 |
| 2010 | Finding Clusters in subspaces of very large, multi-dimensional datasetsabstractWe propose the Multi-resolution Correlation Cluster detection (MrCC), a novel, scalable method to detect correlation clusters able to analyze dimensional data in the range of around 5 to 30 axes. Existing methods typically exhibit super-linear behavior in terms of space or execution time. MrCC employs a novel data structure based on multi-resolution and gains over previous approaches in: (a) it finds clusters that stand out in the data in a statistical sense; (b) it is linear on running time and memory usage regarding number of data points and dimensionality of subspaces where clusters exist; (c) it is linear in memory usage and quasi-linear in running time regarding space dimensionality; and (d) it is accurate, deterministic, robust to noise, does not require stating the number of clusters as input parameter, does not perform distance calculation and is able to detect clusters in subspaces generated by original axes or linear combinations of original axes, including space rotation. We performed experiments on synthetic data ranging from 5 to 30 axes and from 12 k to 250 k points, and MrCC outperformed in time five of the recent and related work, being in average 10 times faster than the competitors that also presented high accuracy results for every tested dataset. Regarding real data, MrCC found clusters at least 9 times faster than the competitors, increasing their accuracy in up to 34 percent. Robson L. F. Cordeiro, Agma J. M. Traina, Christos Faloutsos, Caetano Traina Jr. |
ICDE | 2 |
| 2010 | QMAS: Querying, Mining and Summarization of Multi-modal DatabasesabstractGiven a large collection of images, very few of which have labels, how can we guess the labels of the remaining majority, and how can we spot those images that need brand new labels, different from the existing ones? Current automatic labeling techniques usually scale super linearly with the data size, and/or they fail when only a tiny amount of labeled data is provided. In this paper, we propose QMAS (Querying, Mining And Summarization of Multi-modal Databases), a fast solution to the following problems: (i) low-labor labeling (L3) – given a collection of images, very few of which are labeled with keywords, find the most suitable labels for the remaining ones, and (ii) mining and attention routing – in the same setting, find clusters, the top-NO outlier images, and the top-NR representative images. We report experiments on real satellite images, two large sets (1.5GB and 2.25GB) of proprietary images and a smaller set (17MB) of public images. We show that QMAS scales linearly with the data size, being up to 40 times faster than top competitors (GCap), obtaining better or equal accuracy. In contrast to other methods, QMAS does low-labor labeling (L3), that is, it works even with tiny initial label sets. It also solves both presented problems and spots tiles that potentially require new labels. Robson L. F. Cordeiro, Fan Guo 0006, Donna S. Haverkamp, James H. Horne, Ellen K. Hughes, Gunhee Kim, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
ICDM | 7 |
| 2010 | New DTW-based method to similarity search in sugar cane regions represented by climate and remote sensing time seriesabstractBrazil is an important sugar cane producer, which is the main resource for ethanol production, a renewable source of energy. This agricultural commodity is important to the country economy, becoming fundamental to improve models that assist the crops monitoring process. Vegetation indexes originated from remote sensing images and agrometeorological indexes can be combined to represent sugar cane fields in a regional scale. However, finding different regions with similar patterns to classify or analyze their characteristics is a non-trivial task. Accordingly, this paper presents a method to find similar sugar cane fields represented by series of vegetation and agrometeorological indexes. The proposed method combines a weighted distance function with an algorithm to find similar objects. Results were coincident in the most cases with the classification done by experts, finding regions with similar characteristics of climate and productivity. Consequently, this approach can help in decision making processes by agricultural entrepreneurs. Luciana A. S. Romani, Renata Ribeiro do Valle Gonçalves, Jurandir Zullo Jr., Caetano Traina Jr., Agma J. M. Traina |
IGARSS | 5 |
| 2010 | Metric Data Analysis Enhanced through Temporal VisualizationabstractThe human vision can naturally interpret data in spaces of 2 or 3 dimensions. When data is in higher dimensional spaces, in most cases the visualization is not intuitive. Regarding metric spaces, the interpretation is even harder, since they often do not have a direct spatial representation. However, the need to analyze how metric-represented data evolve over time is pretty common when one needs to understand several phenomena and in decision making processes, as it occurs in medical and agrometeorological applications. This paper presents three interactive techniques to visualize metric data that vary over time. Each one focus on a different way to interpret the temporal information. The first technique shows data evolving in a timeline axis. The second overlaps evolving snapshots of the space showing how the space varies regarding time. The last one does not treat temporal data as a dimension, it is used instead to define the similarity among complex data, employing the new concept of metric-temporal spaces, which seamlessly integrate time and metric data into a single similarity space. Visualization examples with real datasets are presented to show the usefulness of the proposed techniques. Renato Bueno, Humberto Luiz Razente, Daniel S. Kaster, Maria Camila Nardini Barioni, Agma J. M. Traina, Caetano Traina Jr. |
IV | 5 |
| 2010 | Combining Visual Analytics and Content Based Data Retrieval Technology for Efficient Data AnalysisabstractOne of the most useful techniques to help visual data analysis systems is interactive filtering (brushing). However, visualization techniques often suffer from overlap of graphical items and multiple attributes complexity, making visual selection inefficient. In these situations, the benefits of data visualization are not fully observable because the graphical items do not pop up as comprehensive patterns. In this work we propose the use of content-based data retrieval technology combined with visual analytics. The idea is to use the similarity query functionalities provided by metric space systems in order to select regions of the data domain according to user-guidance and interests. After that, the data found in such regions feed multiple visualization workspaces so that the user can inspect the correspondent datasets. Our experiments showed that the methodology can break the visual analysis process into smaller problems (views) and that the views hold the expectations of the analyst according to his/her similarity query selection, improving data perception and analytical possibilities. Our contribution introduces a principle that can be used in all sorts of visualization techniques and systems, this principle can be extended with different kinds of integration visualization-metric-space, and with different metrics, expanding the possibilities of visual data analysis in aspects such as semantics and scalability. José F. Rodrigues Jr., Luciana A. S. Romani, Agma J. M. Traina, Caetano Traina Jr. |
IV | 3 |
| 2010 | Efficient bulk-loading on dynamic metric access methods
Thiago Galbiatti Vespa, Caetano Traina Jr., Agma J. M. Traina |
Inf. Syst. | 3 |
| 2009 | The Onion-Tree: Quick Indexing of Complex Data in the Main Memory
Caio César Mori Carélo, Ives Rene Venturini Pola, Ricardo Rodrigues Ciferri, Agma J. M. Traina, Caetano Traina Jr., Cristina Dutra de Aguiar Ciferri |
ADBIS | 4 |
| 2009 | Unsupervised scaling of multi-descriptor similarity functions for medical image datasetsabstractContent-based search has proven to be a proper complement to textual queries over medical image databases. In many applications, employing multiple image descriptors and combining the respective distance functions using adequate scale factors improves the retrieval accuracy. However, the existing weighting methods are either exhaustive or supervised. In this paper, we present the Fractal-scaled Product Metric, an unsupervised method to determine a scale factor among features in multi-descriptor image similarity assessment based on the fractal theory. The composite distance function obtained is not limited to dimensional image descriptors and enables using scalable indexing structures. Experiments have shown that the proposed method determines near-optimal scale factors for the descriptors involved, and always improves the precision of the results, outperforming the individual descriptors up to 31% on the average precision. Renato Bueno, Daniel S. Kaster, Adriano Arantes Paterlini, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 4 |
| 2009 | Content-based retrieval of medical images: From context to perceptionabstractA challenge in content-based retrieval of image exams is to provide a timely answer that complies to the specialist's expectation. In many situations, when a specialist gets a new image to analyze, having information and knowledge from similar cases can be very helpful. However, the semantic gap between low-level image features and their high level semantics may impair the system acceptability. In this paper we propose a new method where we gather from the physicians the visual patterns they use to recognize anomalies in images and apply this knowledge not only in the preprocessing of the images, but also on building feature extractors based on these visual patterns. Moreover, our approach generates feature vectors with lower dimensionality diminishing the ldquodimensionality curserdquo problem. Experiments using computed tomography lung images show that the proposed method improves the precision of the query results up to 75%, and generates feature vectors up to 94% smaller than traditional feature extraction techniques while keeping the same representative power. This work shows that perception-based feature extraction combined with the image context can be successfully employed to perform similarity queries in medical image databases. Pedro Henrique Bugatti, Marcelo Ponciano-Silva, Agma J. M. Traina, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques |
CBMS | 3 |
| 2009 | Including the perceptual parameter to tune the retrieval ability of pulmonary CBIR systemsabstractThe research on Content-Based Image Retrieval (CBIR) is growing in relevance at a fast pace. Algorithms and tools for CBIR can help decision-making processes, for example allowing the specialist to retrieve cases similar to the one under evaluation. However, the main reservation about using CBIR is the semantic gap, which is the divergence among automatic results and what the user is expecting. We propose the ldquoperceptual parameterrdquo, which allows changing the relationship between the feature extraction algorithms and the distance functions, aimed at finding the best integration of both from the specialist's point of view. This work integrates the three main elements of similarity queries: the extracted features from the images, the distance function employed to quantify the similarity and the similarity perception from the user. These three elements allowed to build the "similarity operators". The experiments performed show that the new perceptual parameter can narrow the semantic gap between what the system retrieves and what the specialist expects. Marcelo Ponciano-Silva, Agma J. M. Traina, Paulo Mazzoncini de Azevedo Marques, Joaquim Cezar Felipe, Caetano Traina Jr. |
CBMS | 2 |
| 2009 | Ranking evaluation functions to improve genetic feature selection in content-based image retrieval of mammogramsabstractThe ranking problem is a crucial task in the information retrieval systems. In this paper, we take advantage of single valued ranking evaluation functions in order to develop a new method of genetic feature selection tailored to improve the accuracy of content-based image retrieval systems. We propose to boost the feature selection ability of the genetic algorithms (GA) by employing an evaluation criteria (fitness function) that relies on order-based ranking evaluation functions. The evaluation criteria are provided by the GA and has been successfully employed as a measure to evaluate the efficacy of content-based image retrieval process, improving up to 22% the precision of the query answers. Experiments on three medical datasets containing breast cancer diagnosis and breast tissue density analysis showed that fitness functions based on ranking evaluation functions occupy an essential role on the algorithms' performance, obtaining results significatively better than other fitness function designs. The experiments also showed that the proposed method obtains results superior than feature selection based on the traditional decision-tree C4.5, naive bayes, support vector machine, 1-nearest neighbor and association rule mining. Sérgio Francisco da Silva, Agma J. M. Traina, Marcela X. Ribeiro, João Batista Neto, Caetano Traina Jr. |
CBMS | 2 |
| 2009 | Time-Aware Similarity Search: A Metric-Temporal Representation for Complex Data
Renato Bueno, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSTD | 3 |
| 2009 | Easing the Dimensionality Curse by Stretching Metric Spaces
Ives Rene Venturini Pola, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 2 |
| 2009 | Supporting content-based image retrieval and computer-aided diagnosis systems with association rule-based techniques
Marcela X. Ribeiro, Pedro Henrique Bugatti, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques, Natalia Abdala Rosa, Agma J. M. Traina |
Data Knowl. Eng. | 6 |
| 2009 | Seamlessly integrating similarity queries in SQLabstractAbstract Modern database applications are increasingly employing database management systems (DBMS) to store multimedia and other complex data. To adequately support the queries required to retrieve these kinds of data, the DBMS need to answer similarity queries. However, the standard structured query language (SQL) does not provide effective support for such queries. This paper proposes an extension to SQL that seamlessly integrates syntactical constructions to express similarity predicates to the existing SQL syntax and describes the implementation of a similarity retrieval engine that allows posing similarity queries using the language extension in a relational DBMS. The engine allows the evaluation of every aspect of the proposed extension, including the data definition language and data manipulation language statements, and employs metric access methods to accelerate the queries. Copyright © 2008 John Wiley & Sons, Ltd. Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr. |
Softw. Pract. Exp. | 3 |
| 2008 | Content-Based Retrieval of Medical Images by Continuous Feature SelectionabstractFeature selection can significantly improve the precision of content-based queries in image databases by removing noisy features or by bursting the most relevant ones. Continuous feature selection techniques assign continuous weights to each feature according to their relevance. In this paper, we propose a supervised method for continuous feature selection. The proposed method applies statistical association rules to find patterns relating low-level image features to high-level knowledge about the images, and it uses the patterns mined to determine the weight of the features. The feature weighting through the statistical association rules reduces the semantic gap that exists between low-level features and the high-level user interpretation of images, improving the precision of the content-based queries. Moreover, the proposed method performs dimensionality reduction of image features avoiding the "dimensionality curse" problem. Experiments show that the proposed method improves the precision of the query results up to 38%, indicating that statistical association rules can be successfully employed to perform continuous feature selection in medical image databases. Pedro Henrique Bugatti, Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 3 |
| 2008 | How to Improve Medical Image Diagnosis through Association Rules: The IDEA MethodabstractIn this paper we present a new method, called IDEA, which employs association rules to assist in medical image diagnosis. IDEA mines association rules, relating visual features with the knowledge gotten from specialists, and employs the associations to suggest possible diagnoses for a given medical image. IDEA incorporates two new algorithms called Omega and ACE. Omega performs simultaneously feature selection and data discretization very efficiently with linear cost on the number of feature values. ACE is a new associative classifier, which has the particular ability of suggesting multiple keywords to compose the diagnosis for a given medical image. The IDEA method has an important characteristic that makes it different from other CAD methods: it suggests multiple diagnosis hypotheses for an image and ranks them based on a measure of quality. The IDEA method was implemented in a prototype (IDEA system) for radiologists evaluate it. The radiologists showed enormous interest in employing the system to aid them in their daily work. The IDEA system was applied to real datasets and the results presented high accuracy (up to 96.7%). The results testify that association rules are well-suited to support the diagnosing task. Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr., Natalia Abdala Rosa, Paulo Mazzoncini de Azevedo Marques |
CBMS | 2 |
| 2008 | A novel optimization approach to efficiently process aggregate similarity queries in metric access methodsabstractA similarity query considers an element as the query center and searches a dataset to find either the elements far up to a bounding radius or the k nearest ones from the query center. Several algorithms have been developed to efficiently execute similarity queries. However, there are queries that require more than one center, which we call Aggregate Similarity Queries. Such queries appear when the user gives multiple desirable examples, and requests data elements that are similar to all of the examples, as in the case of applying relevance feedback. Here we give the first algorithms that can handle aggregate similarity queries on Metric Access Methods (MAM) such as the M-tree and Slim-tree. Our method, which we call Metric Aggregate Similarity Search (MASS) has the following properties: (a) it requires only the triangle inequality property; (b) it guarantees no false-dismissals, as we prove that it lower-bounds the aggregate distance scores; (c) it can work with any MAM; (d) it can handle any number of query centers, which are either scattered all over the space or concentrated on a restricted region. Experiments on both real and synthetic data show that our method scales on both the number of elements and, if the dataset is in a spatial domain, also on its dimensionality. Moreover, it achieves better results than previous related methods. Humberto Luiz Razente, Maria Camila Nardini Barioni, Agma J. M. Traina, Christos Faloutsos, Caetano Traina Jr. |
CIKM | 3 |
| 2008 | A New Approach for Optimization of Dynamic Metric Access Methods Using an Algorithm of Effective Deletion
Renato Bueno, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 3 |
| 2008 | Accelerating k-medoid-based algorithms through metric access methods
Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr. |
J. Syst. Softw. | 3 |
| 2008 | An Association Rule-Based Method to Support Medical Image Diagnosis With EfficiencyabstractIn this paper, we propose a method based on association rule-mining to enhance the diagnosis of medical images (mammograms). It combines low-level features automatically extracted from images and high-level knowledge from specialists to search for patterns. Our method analyzes medical images and automatically generates suggestions of diagnoses employing mining of association rules. The suggestions of diagnosis are used to accelerate the image analysis performed by specialists as well as to provide them an alternative to work on. The proposed method uses two new algorithms, PreSAGe and HiCARe. The PreSAGe algorithm combines, in a single step, feature selection and discretization, and reduces the mining complexity. Experiments performed on PreSAGe show that this algorithm is highly suitable to perform feature selection and discretization in medical images. HiCARe is a new associative classifier. The HiCARe algorithm has an important property that makes it unique: it assigns multiple keywords per image to suggest a diagnosis with high values of accuracy. Our method was applied to real datasets, and the results show high sensitivity (up to 95%) and accuracy (up to 92%), allowing us to claim that the use of association rules is a powerful means to assist in the diagnosing task. Marcela X. Ribeiro, Agma J. M. Traina, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques |
IEEE Trans. Multim. | 2 |
| 2007 | The MM-Tree: A Memory-Based Metric Tree Without Overlap Between Nodes
Ives Rene Venturini Pola, Caetano Traina Jr., Agma J. M. Traina |
ADBIS | 3 |
| 2007 | HEAD: The Human Encephalon Automatic DelimiterabstractIn this paper we present HEAD, the Human Encephalon Automatic Delimiter, a new and efficient method for skull-stripping in T1-weighted MRI that combines an unique histogram analysis with binary mathematical morphology. In our experiments we use real images with highly variable noise ratios and intensity non-uniformity. We evaluate our results based on manually generated true masks and the well known Jaccard metric, achieving accuracy close to 99%. We compare our method with the popular Brain Extractor Surface algorithm (BSE), which in the same experiments achieved less than 95% of accuracy. André G. R. Balan, Agma J. M. Traina, Marcela X. Ribeiro, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr. |
CBMS | 2 |
| 2007 | SuGAR: A Framework to Support Mammogram DiagnosisabstractIn this paper we present a framework based on association-rules to help diagnosis of mammogram abnormalities. Our framework - SuGAR - combines low-level features automatically extracted from images with high-level knowledge gotten from specialists to mine association rules, suggesting possible diagnoses. Our framework is optimized, in the sense that it combines, in a single step, feature selection and discretization, reducing the mining complexity. The framework was applied to real datasets and the results show high sensitivity (up to 95%) and accuracy (up to 92%), allowing us to claim that association rules can effectively aid in the diagnosing task. Marcela X. Ribeiro, Agma J. M. Traina, André G. R. Balan, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques |
CBMS | 2 |
| 2007 | An efficient framework for similarity query optimizationabstractThe increasing volume of multimedia data stored in relational database management systems (RDBMS) demands efficient ways to process similarity queries. Therefore, the query processor should provide mechanisms to express similarity queries, to interpret and translate them into equivalent expression in relational algebra, to evaluate alternative query plans and finally to execute the queries using the best plan found. In this paper, we present an effective framework to interpret, translate, select the best plan and efficiently execute similarity queries over data indexed by metric access methods. Experimental evaluation of the framework shows a reduction of up to 20% in the total time required to answer similarity queries. Mônica Ribeiro Porto Ferreira, Caetano Traina Jr., Agma J. M. Traina |
GIS | 3 |
| 2007 | A Density-Biased Sampling Technique to Improve Cluster Representativeness
Ana Paula Appel, Adriano Arantes Paterlini, Elaine P. M. Sousa, Agma J. M. Traina, Caetano Traina Jr. |
PKDD | 4 |
| 2007 | MAMCost: Global and Local Estimates leading to Robust Cost Estimation of Similarity QueriesabstractThis paper presents an effective cost model to estimate the number of disk accesses (I/O cost) and the number of distance calculations (CPU cost) to process similarity queries over data indexed by metric access methods. Two types of similarity queries were taken into consideration: range and k-nearest neighbor queries. The main point of the cost model is considering not only global parameters of the data set but also the local data distribution. The model takes advantage of the intrinsic dimension of the data set, estimated by its correlation fractal dimension. Experiments were performed on real and synthetic data sets, with different sizes and dimensions, in order to validate the proposed model. They confirmed that the estimations are accurate, within the range achieved by real queries. Gisele Busichia Baioco, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 2 |
| 2007 | Boosting k-Nearest Neighbor Queries Estimating Suitable Query RadiiabstractThis paper proposes novel and effective techniques to estimate a radius to answer k-nearest neighbor queries. The first technique targets datasets where it is possible to learn the distribution about the pairwise distances between the elements, generating a global estimation that applies to the whole dataset. The second technique targets datasets where the first technique cannot be employed, generating estimations that depend on where the query center is located. The proposed k-NNF() algorithm combines both techniques, achieving remarkable speedups. Experiments performed on both real and synthetic datasets have shown that the proposed algorithm can accelerate k-NN queries more than 26 times compared with the incremental algorithm and spends half of the total time compared with the traditional k-NN() algorithms. Marcos R. Vieira, Caetano Traina Jr., Agma J. M. Traina, Adriano S. Arantes, Christos Faloutsos |
SSDBM | 3 |
| 2007 | A fast and effective method to find correlations among attributes in databases
Elaine P. M. Sousa, Caetano Traina Jr., Agma J. M. Traina, Leejay Wu, Christos Faloutsos |
Data Min. Knowl. Discov. | 3 |
| 2007 | Genetic algorithms for approximate similarity queries
Renato Bueno, Agma J. M. Traina, Caetano Traina Jr. |
Data Knowl. Eng. | 2 |
| 2007 | Investigating the potential of art neural network models for indexing and information retrievalabstractDatabase management systems are very sophisticated, efficient, and fast in information retrieval tasks involving traditional data sets such as numbers, strings, and so on, but many limitations become evident when the data are more complex, that is, high or nondimensional data. Considering some existing problems in information retrieval processes, this work proposes a hybrid system that combines a model of the ART family neural network, ART2-A, with the Slim-Tree data structure, which is a metric access method. This approach is an alternative to perform clustering on data in an intelligent way so that the data can be recovered from the corresponding Slim-Tree. The proposed hybrid system is able to perform range and k-nearest neighbor queries, which is not an inherent characteristic in implementations involving artificial neural networks. Furthermore, experimental results showed that the performance of the hybrid system was better than the performance of Slim-Tree. © 2007 Wiley Periodicals, Inc. Int J Int Syst 22: 319–336, 2007. Roseli A. Francelin Romero, José F. Vicentini, Patrícia R. Oliveira 0001, Agma J. M. Traina |
Int. J. Intell. Syst. | 4 |
| 2007 | The Omni-family of all-purpose access methods: a simple and effective way to make similarity search more efficient
Caetano Traina Jr., Roberto F. Santos Filho, Agma J. M. Traina, Marcos R. Vieira, Christos Faloutsos |
VLDB J. | 3 |
| 2006 | Distance Functions Association for Content-Based Image Retrieval using Multiple Comparison CriteriaabstractThe comparison operators available in traditional Database Management Systems (DBMS) are not adequate to handle complex data such as images, rather comparing them using similarity operators is the option of choice. Similarity operators need a way to measure the similarity between pairs of objects. Although there are many interesting works dealing with similarity queries and functions to measure similarity, they all rely on a single similarity function that must be applicable over the whole dataset. However, images from medical exams often require several ways to measure similarity, depending on many factors, such as the particular pathological condition being searched, or the existence of specific clinical condition revealed in the images compared. Therefore, the ability to handle several ways to compare images by similarity is important in medical software handling images. This work develop a technique to allow several similarity functions to be combined when indexing a large set of images, allowing queries to probe the dataset regarding distinct comparison criteria. This technique also allows a flexible way to pose queries supporting fast retrieval of the answers. Ives Rene Venturini Pola, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2006 | Statistical Association Rules and Relevance Feedback: Powerful Allies to Improve the Retrieval of Medical ImagesabstractThis work aims at developing an efficient support to improve the precision of medical image retrieval by content, introducing an approach that combines techniques of statistical association rule mining and relevance feedback. Low level features of shape and texture are extracted from images. Statistical association rules are used to select the most relevant features to discriminate the images, reducing the size of the feature vectors and eliminating noisy features that influence negatively the query results, making the whole process more efficient. Additionally, our approach uses a new relevance feedback technique to overcome the semantic gap that exists between low level features and the high level user interpretation of images. Experiments show that the combination of statistical association rule mining and the relevance feedback technique proposed here improve the precision of the query results up to 100%. Marcela X. Ribeiro, Joselene Marques, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 3 |
| 2006 | Fighting the Semantic Gap on CBIR Systems through New Relevance Feedback TechniquesabstractThis paper introduces two novel relevance feedback techniques that integrate a new way to implement the query center movement with a suitable weighting on the similarity function. These techniques integrated to a content-based image retrieval (CBIR) system, improves the precision of the results when using texture features up to 42%, and employing at most 5 iterations. Thus, the user satisfaction with the system is increased as our experiments demonstrated. Besides being effective, the new RF techniques are very fast as they take less than one second to reprocess the queries at each iteration. The experiments also show that with three iterations the users are satisfied with the query results, and the major gain in precision happens in the first iteration, achieving improvements of up to 30%, what lessens the user efforts and anxiety. Agma J. M. Traina, Joselene Marques, Caetano Traina Jr. |
CBMS | 1 |
| 2006 | Efficient processing of complex similarity queries in RDBMS through query rewritingabstractMultimedia and complex data are usually queried by similarity predicates. Whereas there are many works dealing with algorithms to answer basic similarity predicates, there are not generic algorithms able to efficiently handle similarity complex queries combining several basic similarity predicates. In this work we propose a simple and effective set of algorithms that can be combined to answer complex similarity queries, and a set of algebraic rules useful to rewrite similarity query expressions into an adequate format for those algorithms. Those rules and algorithms allow relational database management systems to turn complex queries into efficient query execution plans. We present experiments that highlight interesting scenarios. They show that the proposed algorithms are orders of magnitude faster than the traditional similarity algorithms. Moreover, they are linearly scalable considering the database size. Caetano Traina Jr., Agma J. M. Traina, Marcos R. Vieira, Adriano S. Arantes, Christos Faloutsos |
CIKM | 2 |
| 2006 | SuperGraph VisualizationabstractGiven a large social or computer network, how can we visualize it, find patterns, outliers, communities? Although several graph visualization tools exist, they cannot handle large graphs with hundred thousand nodes and possibly million edges. Such graphs bring two challenges: interactive visualization demands prohibitive processing power and, even if we could interactively update the visualization, the user would be overwhelmed by the excessive number of graphical items. To cope with this problem, we propose a formal innovation on the use of graph hierarchies that leads to GMine system. GMine promotes scalability using a hierarchy of graph partitions, promotes concomitant presentation for the graph hierarchy and for the original graph, and extends analytical possibilities with the integration of the graph partitions in an interactive environment. José F. Rodrigues Jr., Agma J. M. Traina, Christos Faloutsos, Caetano Traina Jr. |
ISM | 2 |
| 2006 | Reviewing Data Visualization: an Analytical Taxonomical StudyabstractThis paper presents an analytical taxonomy that can suitably describe, rather than simply classify, techniques for data presentation. Unlike previous works, we do not consider particular aspects of visualization techniques, but their mechanisms and foundational vision perception. Instead of just adjusting visualization research to a classification system, our aim is to better understand its process. For doing so, we depart from elementary concepts to reach a model that can describe how visualization techniques work and how they convey meaning José F. Rodrigues Jr., Agma J. M. Traina, Maria Cristina Ferreira de Oliveira, Caetano Traina Jr. |
IV | 2 |
| 2006 | Automatic mining of fruit fly embryo imagesabstractWe present FEMine, an automatic system for image-based gene expression analysis. We perform experiments on the largest publicly available collection of Drosophila ISH (in situ hybridization) images, showing that our FEMine system achieves excellent performance in classification, clustering, and content-based image retrieval. The major innovation of FEMine is the use of automatically discovered latent spatial "themes" of gene expressions, LGEs, in the whole-embryo context, as opposed to patterns in nearly disjoint portions of an embryo proposed in previous methods. Jia-Yu Pan, André G. R. Balan, Eric P. Xing, Agma J. M. Traina, Christos Faloutsos |
KDD | 4 |
| 2006 | SIREN: A Similarity Retrieval Engine for Complex Data
Maria Camila Nardini Barioni, Humberto Luiz Razente, Agma J. M. Traina, Caetano Traina Jr. |
VLDB | 3 |
| 2006 | GMine: A System for Scalable, Interactive Graph Visualization and Mining
José F. Rodrigues Jr., Hanghang Tong, Agma J. M. Traina, Christos Faloutsos, Jure Leskovec |
VLDB | 3 |
| 2005 | Enhanced visual evaluation of feature extractors for image miningabstractSummary form only given. This paper introduces a novel approach to evaluate, timely and effectively, the suitability of new image feature extraction techniques concerning similarity queries using CBIR systems. The proposed approach is based on two measurements derived from spatial properties intuitively and naturally perceived in spatial domains, and that can also be verified in multidimensional spaces. To bear out our proposal, we show that the insights obtained by the proposed measurements comply with the well-known analysis methods based on the precision and recall approach. José F. Rodrigues Jr., Agma J. M. Traina, Caetano Traina Jr. |
AICCSA | 2 |
| 2005 | Fractal Analysis of Image Textures for Indexing and Retrieval by ContentabstractThis paper proposes the use of fractal analysis as a means to discriminate textured segmented regions of medical images. We show that the use of the fractals can boost the representation level of traditional image features allowing high rates of precision when answering similarity queries over images employing a variance weighted Manhattan distance. The cost to compute the fractal measurements is linear on the image size, what makes their use a suitable choice for large sets of images. André G. R. Balan, Agma J. M. Traina, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques |
CBMS | 2 |
| 2005 | A Low-cost Approach for Effective Shape-based Retrieval and Classification of Medical ImagesabstractThis work aims at developing an efficient support for retrieval and classification of medical images, introducing an approach that comprises techniques of image processing, data mining and fractal theory, leading to an effective and direct way to compare images. A method of feature extraction and comparison is proposed, which uses Zernike moments for invariant pattern recognition as shape features of images' regions of interest. A new algorithm that generates statistical-based association rules is used to identify representative features that discriminate the disease classes of images. In order to minimize the computational effort, another new algorithm, based on fractal theory, is applied to reduce the dimensionality of the representative feature space. In essence, the proposed method determines the smallest set of relevant features that can properly represent images without loss of precision. In addition, the method discards the need of image segmentation, leading to a simple but effective way to make image retrieval by content. Experiments executing k-nearest neighbor queries on medical images reveal that the process is robust and suitable to perform retrieval combined with classification of this kind of images. Joaquim Cezar Felipe, Jonatas B. Olioti, Agma J. M. Traina, Marcela X. Ribeiro, Elaine P. M. Sousa, Caetano Traina Jr. |
ISM | 3 |
| 2005 | Global Warp Metric Distance: Boosting Content-based Image Retrieval through HistogramsabstractThis work presents a new distance function - the global warp metric distance - to compare histograms used as a feature to index image databases in content based image retrieval environments. The metric histogram represents a compact, but efficient alternative to the use of traditional gray level histograms to represent images. The global warp metric distance (GWMD) enhances the comparison between histograms, replacing the rigid bin to bin evaluation by the warp method, which allows a local "adjustment" of one histogram to the other during the distance calculation, introducing a global matching of the curves. Besides this, GWMD applies a set of geometric global features of histograms to determine the final distance. Results on similarity retrieval in medical images demonstrate the superiority of the proposed approach in analyzing image sets that present brightness and contrast disparities: it reduces the amount of both false positive and false negative retrievals. Moreover, these results comply with similarity evaluations performed by domain specialists. Joaquim Cezar Felipe, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 2 |
| 2004 | Content-based Image Retrieval Using Approximate Shape of ObjectsabstractThis paper presents a new approach to retrieve images by content using a composition of relevant features regarding texture, shape and brightness distribution. The first step of the method is a segmentation process based on Markov random fields, which can be done automatically, having as parameter the number of desired classes. The regions obtained in the segmentation guide the extraction of measures from the original image producing a 30-dimensional feature vector used in the image retrieval. The experiments showed that the feature vector has high discrimination power and the time for retrieval operations are only fractions of seconds. Agma J. M. Traina, André G. R. Balan, Luis M. Bortolotti, Caetano Traina Jr. |
CBMS | 1 |
| 2004 | Including Conditional Operators in Content-Based Image Retrieval in Large Sets of Medical ExamsabstractContent-based image retrieval systems (CBIR) aim at helping in searching large image collections to find those more likely to answer query conditions based on the information represented in the images. To speed up the search process, selected features are extracted from each image when they are stored in the database, so each one is represented by a feature vector. Subsequent image searching operations are performed using the feature vectors in place of the images. The feature extraction algorithms have important issues in CBIR due to the large semantic gap between the low-level features extracted as compared to the high-level, semantic, results expected by the users. A way to approach a semantic analysis of an image, as performed by humans, is to employ a large number of analyzers whose results are processed by a set of rules based on if-then clauses. In this paper we create a framework to define image processing and feature extraction algorithms for CBIR systems as components of a data flow architecture, including an analysis mechanism to interpret images. It is based on decision components that check the results of previously executed feature extractors to choose from a set of configurable execution paths that leads to the creation of the feature vector of each image. In the paper we will describe a real system that has been implemented based on these concepts, which is being used as a teaching tool in a school hospital. Caetano Traina Jr., Agma J. M. Traina, Josiel Maimoni de Figueiredo |
CBMS | 2 |
| 2004 | Knowledge Extraction using Visualization of Hemoglobin Parameters to Identify ThalassemiaabstractThe analysis of large amounts of data is better performed by humans when represented in a graphical format. Therefore, a new research area called the visual data mining is being developed endeavoring to use the number crunching power of computers to prepare data for visualization, allied to the ability of humans to interpret data presented graphically. This work presents the results of applying a visual data mining tool, called FastMapDB to detect the behavioral pattern exhibited by a dataset of clinical information about hemoglobinopathies known as thalassemia. FastMapDB is a visual data mining tool that get tabular data stored in a relational database such as dates, numbers and texts, and by considering them as points in a multidimensional space, maps them to a three-dimensional space. The intuitive three-dimensional representation of objects enables a data analyst to "see" the behavior of the characteristics from abnormal forms of hemoglobin, highlighting the differences when compared to data from a group without alteration. Carlos Roberto Valêncio, Mauricio N. Tronco, Ana C. Bonini-Domingos, Cláudia R. Bonini-Domingos, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 6 |
| 2003 | Retrieval by Content of Medical Images Using Texture for Tissue IdentificationabstractThis work aims at supporting the retrieval and indexing of medical images by extracting and organizing intrinsic features of them, more specifically texture attributes from images. A tool for obtaining the relevant textures was implemented This tool retrieves and classifies images using the extracted values, and allows the user to issue similarity queries. The application of the proposed method on images has given encouraging results that motivate to apply the method as a basis to more experiments, at diversified contexts. The accuracy degree obtained from the precision and recall plots was always over 90% for queries asking for similar images for up to 20% of the database. Joaquim Cezar Felipe, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2003 | MultiWaveMed: A System for Medical Image Retrieval through Wavelets TransformationsabstractThis paper presents the MultiWaveMed system, which is a new software allowing to index and retrieve medical images through the comparison of their texture features. The features are extracted by wavelet transforms, and are organized in feature vectors. The system extracts the image texture features, computes the distance between the query image to all images in the database, through the comparison of their features, and retrieve de n most similar images regarding this kind of feature. The proposed system has implemented both Daubechies and Gabor wavelets. The feature vectors extracted from the images are used to organize the images through access methods, which are the basis to perform the query-by-content operations over the images. The focus of this paper is to show the utility of the wavelet transforms on medical image characterization and their suitability for image indexing and retrieval. Agma J. M. Traina, César A. B. Castañón, Caetano Traina Jr. |
CBMS | 1 |
| 2003 | Integrating Images to Patient Electronic Medical Records through Content-Based Retrieval TechniquesabstractThis paper presents the SRIS-HC-an Patient Electronic Medical Record System with support for Content-based Image Retrieval, developed aiming at demonstrating the benefits of having the ability of similarity retrieval over image datasets based on their contents, at the Clinical Hospital of the Medical School of Ribeirao Preto of the University of Sao Paulo at Ribeirao Preto-Brazil (the HCFMRP/USP). This ability is an additional resource developed over a PACS system, intending to enable the improvement of medical diagnosis by images, as well as to provide a basis to perform analysis of similar medical cases and bibliographic research over them. The SRIS-HC was developed on top of the Radiology Information System (RIS) of the Radio-diagnosis Laboratory of the HCFMRP/USP, which is in turn at the core of the Electronic Patient Record System of the hospital. Agma J. M. Traina, Natalia Abdala Rosa, Caetano Traina Jr. |
CBMS | 1 |
| 2003 | Efficient Content-Based Image Retrieval through Metric Histograms
Agma J. M. Traina, Caetano Traina Jr., Josiane Maria Bueno, Fabio Jun Takada Chino, Paulo Mazzoncini de Azevedo Marques |
World Wide Web | 1 |
| 2002 | Extending Relational atabases to Support Content-based Retrieval of Medical ImagesabstractThis paper shows how to support images in a relational database, so it can fulfill the requirements to be used as the storage mechanism of a PACS. This support includes the ability to answer similarity queries based on the image content, providing fast image retrieval based on indexing structures. The main concept allowing this support is the definition of distance functions based on features, which are extracted from the images as they are stored in the database. An extension to SQL enables the construction of an interpreter that intercepts the extended commands and translates them into standard SQL, allowing one to take advantage of any relational database server. We describe experiments made with a prototype implemented using these concepts, which allowed answering queries up to 20 times faster than using existing relational servers alone. Myrian R. B. Araujo, Caetano Traina Jr., Agma J. M. Traina, Josiane Maria Bueno, Humberto Luiz Razente |
CBMS | 3 |
| 2002 | How to Add Content-based Image Retrieval Capability in a PACSabstractThis paper presents a new picture archiving and communication system (PACS), called cbPACS (content-based PACS), which has content-based image retrieval resources. cbPACS answers similarity (range and nearest-neighbor) queries, taking advantage of a metric access method embedded into the image database manager. The images are compared via their features, which are extracted by an image processing system module. The system works on features based on the color distribution of the images through normalized histograms as well as metric histograms. Metric histograms are invariant with regard to scale, translation and rotation of images and also to brightness transformations. cbPACS is prepared to integrate new image features, based on the texture and shape of the main objects in the image. Josiane Maria Bueno, Fabio Jun Takada Chino, Agma J. M. Traina, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques |
CBMS | 3 |
| 2002 | How to improve the pruning ability of dynamic metric access methodsabstractComplex data retrieval is accelerated using index structures, which organize the data in order to prune comparisons between data during queries. In metric spaces, comparison operations can be specially expensive, so the pruning ability of indexing methods turns out to be specially meaningful. This paper shows how to measure the pruning power of metric access methods, and defines a new measurement, called "prunability," which indicates how well a pruning technique carries out the task of cutting down distance calculations at each tree level. It also presents a new dynamic access method, aiming to minimize the number of distance calculations required to answer similarity queries. We show that this novel structure is up to 3 times faster and requires less than 25% distance calculations to answer similarity queries, as compared to existing methods. This gain in performance is achieved by taking advantage of a set of global representatives. Although our technique uses multiple representatives, the index structure still remains dynamic and balanced. Caetano Traina Jr., Agma J. M. Traina, Roberto F. Santos Filho, Christos Faloutsos |
CIKM | 2 |
| 2002 | Fast Indexing and Visualization of Metric Data Sets using Slim-TreesabstractMany recent database applications need to deal with similarity queries. For such applications, it is important to measure the similarity between two objects using the distance between them. Focusing on this problem, this paper proposes the slim-tree, a new dynamic tree for organizing metric data sets in pages of fixed size. The slim-tree uses the triangle inequality to prune the distance calculations that are needed to answer similarity queries over objects in metric spaces. The proposed insertion algorithm uses new policies to select the nodes where incoming objects are stored. When a node overflows, the slim-tree uses a minimal spanning tree to help with the splitting. The new insertion algorithm leads to a tree with high storage utilization and improved query performance. The slim-tree is a metric access method that tackles the problem of overlaps between nodes in metric spaces and that allows one to minimize the overlap. The proposed "fat-factor" is a way to quantify whether a given tree can be improved and also to compare two trees. We show how to use the fat-factor to achieve accurate estimates of the search performance and also how to improve the performance of a metric tree through the proposed "slim-down" algorithm. This paper also presents a new tool in the slim-tree's arsenal of resources, aimed at visualizing it. Visualization is a powerful tool for interactive data mining and for the visual tracking of the behavior of a tree under updates. Finally, we present a formula to estimate the number of disk accesses in range queries. Results from experiments with real and synthetic data sets show that the new slim-tree algorithms lead to performance improvements. These results show that the slim-tree outperforms the M-tree by up to 200% for range queries. For insertion and splitting, the minimal-spanning-tree-based algorithm achieves up to 40 times faster insertions. We observed improvements of up to 40% in range queries after applying the slim-down algorithm. Caetano Traina Jr., Agma J. M. Traina, Christos Faloutsos, Bernhard Seeger |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | Similarity Search without Tears: The OMNI Family of All-purpose Access MethodsabstractDesigning a new access method inside a commercial DBMS is cumbersome and expensive. We propose a family of metric access methods that are fast and easy to implement on top of existing access methods, such as sequential scan, R-trees and Slim-trees. The idea is to elect a set of objects as foci, and gauge all other objects with their distances from this set. We show how to define the foci set cardinality, how to choose appropriate foci, and how to perform range and nearest-neighbor queries using them, without false dismissals. The foci increase the pruning of distance calculations during the query processing. Furthermore we index the distances from each object to the foci to reduce even triangular inequality comparisons. Experiments on real and synthetic datasets show that our methods match or outperform existing methods. They are up to 10 times faster, and perform up to 10 times fewer distance calculations and disk accesses. In addition, it scales up well, exhibiting sub-linear performance with growing database size. Roberto F. Santos Filho, Agma J. M. Traina, Caetano Traina Jr., Christos Faloutsos |
ICDE | 2 |
| 2001 | Tri-plots: scalable tools for multidimensional data miningabstractWe focus on the problem of finding patterns across two large, multidimensional datasets. For example, given feature vectors of healthy and of non-healthy patients, we want to answer the following questions: Are the two clouds of points separable? What is the smallest/largest pair-wise distance across the two datasets? Which of the two clouds does a new point (feature vector) come from?We propose a new tool, the tri-plot, and its generalization, the pq-plot, which help us answer the above questions. We provide a set of rules on how to interpret a tri-plot, and we apply these rules on synthetic and real datasets. We also show how to use our tool for classification, when traditional methods (nearest neighbor, classification trees) may fail. Agma J. M. Traina, Caetano Traina Jr., Spiros Papadimitriou, Christos Faloutsos |
KDD | 1 |
| 2000 | Slim-Trees: High Performance Metric Trees Minimizing Overlap Between Nodes
Caetano Traina Jr., Agma J. M. Traina, Bernhard Seeger, Christos Faloutsos |
EDBT | 2 |
| 2000 | Distance Exponent: A New Concept for Selectivity Estimation in Metric TreesabstractThis paper discusses the problem of selectivity estimation for range queries in metric datasets, which include vector, or dimensional, datasets as a special case. The main contribution of this paper is that, surprisingly, many different real datasets follow a law. From this observation we derive an analysis for the distribution of metric datasets. This is the first analysis of distributions for real metric datasets.We called the exponent of our power law as distance exponent. We show that it plays a relevant role for the analysis of real, metric datasets. Specifically, we show (a) how to exploit the exponent to derive formulas for selectivity estimation of range queries and (b) how to compute it quickly from a metric index tree.We performed several experiments on many real datasets (road intersections of U.S. counties, vectors characteristics extracted from face matching systems, sets of words, matrixes) and synthetic datasets (Sierpinsky triangle, a 2-dimensional uniform distribution and a 2-dimensional line). Our selectivity estimation formulas are accurate, within relative error from 4% to 17%, and always within one standard deviation from the analytical results. Moreover, we present also a quick algorithm to estimate the distance exponent, which gives good accuracy and saves orders of magnitude in computation time. Caetano Traina Jr., Agma J. M. Traina, Christos Faloutsos |
ICDE | 2 |
| 2000 | Spatial Join Selectivity Using Power LawsabstractWe discovered a surprising law governing the spatial join selectivity across two sets of points. An example of such a spatial join is “find the libraries that are within 10 miles of schools”. Our law dictates that the number of such qualifying pairs follows a power law, whose exponent we call “pair-count exponent” (PC). We show that this law also holds for self-spatial-joins (“find schools within 5 miles of other schools”) in addition to the general case that the two point-sets are distinct. Our law holds for many real datasets, including diverse environments (geographic datasets, feature vectors from biology data, galaxy data from astronomy). Christos Faloutsos, Bernhard Seeger, Agma J. M. Traina, Caetano Traina Jr. |
SIGMOD Conference | 3 |
| 1997 | 3D reconstruction of magnetic resonance imaging using largely spaced slicesabstractThis paper presents a full reconstruction process of magnetic resonance images. The first step is to bring the acquired data from the frequency domain, using a fast Fourier transform algorithm. A tomographic image interpolation is then used to transform a sequence of tomographic slices in an isotropic volume data set, a process also called 3D reconstruction. This work describes a method the interpolation stage of which is based on a previous matching stage using Delaunay triangulation. Agma J. M. Traina, Afonso H. M. A. Prado, Josiane Maria Bueno |
CBMS | 1 |