Mirela Teixeira Cazzolato

dblp:135/8726 · also Mirela T. Cazzolato · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-4364-010XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 13 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Modeling Clinical Data with Attention: A Knowledge Graph Approach with CliniKG
abstract
Given a set of historical, unlabeled patient records, how can we model the relationships between various concepts related to health conditions and treatments? Modeling health data as knowledge graphs can aid in better understanding and mining recurrent and abnormal patterns within massive amounts of records. However, for generic data modeling, the concepts and relationships of the data are often defined manually, which can be laborious and prone to human errors. The lack of structure in textual reports makes modeling challenging, as there is typically no standard in information, terminology, or other elements. This work proposes CliniKG for utilizing Large Language Models (LLMs) to structure patient records, enabling the modeling of health data concepts. First, we employ LLMs with zero-, or few-shot learning to define the graph's relationships of pairwise concepts. Based on the resulting modeling, we extract meaningful features from nodes and define their semantics automatically. With data visualizations, CliniKG highlights key findings, such as recurrent or rare relationships. The experimental evaluation shows CliniKG in action using a real dataset from a public hospital in Brazil. We performed a qualitative analysis with 40 domain experts to evaluate medium-sized LLMs and prompt con-figurations. The study reveals interesting patterns automatically identified within the data. CliniKG exhibits linear performance relative to the number of nodes, and the visual tools can assist specialists in monitoring patients' conditions.
Eduardo Moura, Rafael C. G. Conrado, Leonardo de Oliveira Campos, Mauro M. Olivatto, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina, Mirela Teixeira Cazzolato
CBMS8
2024 T-NET: Weakly Supervised Graph Learning for Combatting Human Trafficking
abstract
Human trafficking (HT) for forced sexual exploitation, often described as modern-day slavery, is a pervasive problem that affects millions of people worldwide. Perpetrators of this crime post advertisements (ads) on behalf of their victims on adult service websites (ASW). These websites typically contain hundreds of thousands of ads including those posted by independent escorts, massage parlor agencies and spammers (fake ads). Detecting suspicious activity in these ads is difficult and developing data-driven methods is challenging due to the hard-to-label, complex and sensitive nature of the data. In this paper, we propose T-Net, which unlike previous solutions, formulates this problem as weakly supervised classification. Since it takes several months to years to investigate a case and obtain a single definitive label, we design domain-specific signals or indicators that provide weak labels. T-Net also looks into connections between ads and models the problem as a graph learning task instead of classifying ads independently. We show that T-Net outperforms all baselines on a real-world dataset of ads by 7% average weighted F1 score. Given that this data contains personally identifiable information, we also present a realistic data generator and provide the first publicly available dataset in this domain which may be leveraged by the wider research community.
Pratheeksha Nair, Javin Liu, Catalina Vajiac, Andreas M. Olligschlaeger, Polo Chau, Mirela Teixeira Cazzolato, Cara Jones, Christos Faloutsos, Reihaneh Rabbany
AAAI6
2024 A symptom-based community-weighted similarity approach for inpatient health condition monitoring
abstract
Given a patient’s series of exams conducted over time, how can we identify cases with similar abnormalities or symptoms? Hospitals and medical facilities continuously monitor patients through periodic exams, a crucial practice for assessing their current condition and potential progression, thereby supporting decision-making. However, similarity-based searches often consider several exams of a patient, most times overlooking the temporal aspect, which is crucial for patient monitoring. In this paper, we present: (1) a novel similarity search framework that identifies similar cases based on symptoms while considering the temporal evolution of the patients’ conditions; and (2) a novel similarity function, called GCWei function, which is built upon the traditional Levenshtein similarity and improves the quality of the search by penalizing the similarity between non-related sets of symptoms. To identify relations, GCWei relies on well-established graph community detection procedures using all patients’ historical data. By combining (1) and (2), we obtain a search approach called GCWei-based search, which efficiently retrieves similar cases with similar developments and thus gives the specialist a broader view of the patient’s condition based on past cases of other patients. To demonstrate the value of our approach, we evaluate it both quantitatively and qualitatively using the recent and publicly available MIMIC-IV database.
Jean R. Ponciano, Mirela Teixeira Cazzolato, Marco A. Gutierrez 0001, Caetano Traina Jr., Agma J. M. Traina
CBMS2
2023 TgrApp: Anomaly Detection and Visualization of Large-Scale Call Graphs
abstract
Given a million-scale dataset of who-calls-whom data containing imperfect labels, how can we detect existing and new fraud patterns? We propose TgrApp, which extracts carefully designed features and provides visualizations to assist analysts in spotting fraudsters and suspicious behavior. Our TgrApp method has the following properties: (a) Scalable, as it is linear on the input size; and (b) Effective, as it allows natural interaction with human analysts, and is applicable in both supervised and unsupervised settings.
Mirela Teixeira Cazzolato, Saranya Vijayakumar, Namyong Park 0001, Meng-Chieh Lee, Polo Chau, Pedro Fidalgo, Bruno Lages, Agma J. M. Traina, Christos Faloutsos
AAAI1
2023 Exploratory Data Analysis in Electronic Health Records Graphs: Intuitive Features and Visualization Tools
abstract
Given a large, unlabeled set of Electronic Health Records (EHRs) acquired from multiple hospitals, how can we analyze the available entities and identify relationships in the data? Also, how can we perform Exploratory Data Analysis (EDA) over such EHR data? Many medical institutions generate EHRs as tabular data with entities and attributes in common. However, due to a large number of records, attributes, and high cardinality, exploring the different datasets and finding patterns and insights become laborious and prone to errors. In this work, we propose GraF- Eda for EDA over EHR data from different institutions. GraF-EDA models EHRs as time-evolving graphs, allowing the interoperability of such data into a single representation. We extract meaningful features from the graph nodes and provide intuitive visualizations to improve data explainability. We evaluate GraF-EDA with four COVID-19 datasets from hospitals of the São Paulo state, Brazil, resulting in million-scale graphs. Our method identified correlations, similarities and dissimilarities among medical treatments, exams, clinics, and outcomes. With the visual tools provided by GraF-EDA, we were able to spot cases of interest and check more details about them. Our results indicate that GraF-EDA is a fast, effective, open-sourced tool for EDA of EHRs from multiple institutions.
Mirela Teixeira Cazzolato, Marco A. Gutierrez 0001, Caetano Traina Jr., Christos Faloutsos, Agma J. M. Traina
CBMS1
2023 CallMine: Fraud Detection and Visualization of Million-Scale Call Graphs
abstract
Given a million-scale dataset of who-calls-whom data containing imperfect labels, how can we detect existing and new fraud patterns? We propose CallMine, with carefully designed features and visualizations. Our CallMine method has the following properties: (a) Scalable, being linear on the input size, handling about 35 million records in around one hour on a stock laptop; (b) Effective, allowing natural interaction with human analysts; (c) Flexible, being applicable in both supervised and unsupervised settings; (d) Automatic, requiring no user-defined parameters.
Mirela Teixeira Cazzolato, Saranya Vijayakumar, Meng-Chieh Lee, Catalina Vajiac, Namyong Park 0001, Pedro Fidalgo, Agma J. M. Traina, Christos Faloutsos
CIKM1
2022 TgraphSpot: Fast and Effective Anomaly Detection for Time-Evolving Graphs
abstract
Given a large, time-evolving graph of who-calls-whom-when, how can we help analysts find anomalies and fraudsters? How can we explain our decisions? We provide TgraphSpot, which carefully extracts features that are often related to fraud; and which provides informative, interactive plots that help analysts zoom down to the few strange nodes. We present the architecture and design decisions of TgraphSpot. Thanks to our careful feature-extraction algorithms, it scales linearly, taking 2.5 hours on a stock laptop, to process 29 million phone calls. More importantly, when applied on a real dataset of millions of phone calls, it discovered suspicious nodes; experts confirmed that those nodes are fraudsters that had been undetected so far.
Mirela Teixeira Cazzolato, Saranya Vijayakumar, Namyong Park 0001, Meng-Chieh Lee, Pedro Fidalgo, Bruno Lages, Agma J. M. Traina, Christos Faloutsos
IEEE Big Data1
2022 Analysis of vertebrae without fracture on spine MRI to assess bone fragility: A Comparison of Traditional Machine Learning and Deep Learning
abstract
Bone mineral density (BMD) is the international standard for evaluating osteoporosis/osteopenia. The success rate of BMD alone in estimating the risk of vertebral fragility fracture (VFF) is approximately 50%, making BMD far from ideal in predicting VFF. In addition, whether or not a patient has been diagnosed with osteoporosis or osteopenia, he or she may suffer a VFF. For this reason, we conducted an extensive empirical study to assess VFFs in postmenopausal women. We considered a representative dataset of 94 T1- and T2-weighted routine spine MRI (with osteopenia or osteoporosis), split into 2,400 samples (slices). Comparing the classification results of machine learning and deep learning (DL) techniques showed that DL generally achieved better results at the cost of higher computational power and hard explainability. ResNet achieved the best results in discriminating patients from groups with and without VFFs with 83% accuracy and 90% AUC (with a confidence interval of 99%). Our results represent a significant step toward prospective and longitudinal studies investigating methods to achieve higher accuracy in predicting VFFs based on spine MRI features of vertebrae without fracture.
Jonathan S. Ramos, Erikson Júlio De Aguiar, Ivar Vargas Belizario, Márcus V. L. Costa, Jamilly G. Maciel, Mirela Teixeira Cazzolato, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Agma J. M. Traina
CBMS6
2022 Establishing trajectories of moving objects without identities: The intricacies of cell tracking and a solution
abstract
Storing, querying, predicting, and interpolating trajectories of moving objects is a topic which the database community has studied for decades. We study a new variant of this problem in this article: We deal with a set of moving objects which do not have an identity, i.e., one does not know whether an object is identical to one observed earlier at another position. Our use case is a stream of images of cells of developing embryos. There exist so-called tracking tools. They match cells in such image sequences, to build trajectory vectors. However, these trackers have certain weaknesses, including counter-intuitive parameters and the expectation of users manually correcting trajectories. In this paper, we propose fully automatic tracking algorithms. They rely on space partitioning heuristics to match cells. This gives way to much cheaper data-analysis pipelines, as we will explain. We also propose two algorithms predicting the next positions of cells, given earlier ones. Experiments over 12 datasets show that our new approaches reduce the execution time by up to 7.8 times for tracking and 6.2 times for prediction. Prediction quality increases by up to 5.6% over the best tracker. • Cells can be modeled as moving objects without identity that move under uncertainty. • Predictors establish cell motion accurately based on observed cell positions. • Cell prediction avoids computationally costly steps of the tracking pipeline.
Mirela Teixeira Cazzolato, Agma J. M. Traina, Klemens Böhm
Inf. Syst.1
2021 BEAUT: a radiomic approach to identify potential lumbar fractures in magnetic resonance imaging
abstract
Bone densitometry (DEXA) is the international reference standard to evaluate Bone Mineral Density (BMD) and diagnose osteoporosis. However, DEXA is far from ideal when used to predict fragility fractures, which are strongly related to morbidity and mortality. According to the literature, spine MRI texture features correlate well with DEXA measurements. For this reason, we conducted an extensive empirical study aimed at assessing fragility fractures secondary to osteoporosis. To perform the evaluations, we developed a radiomic-based approach called BEAUT (BonE Analysis Using Texture). We performed experiments on a meaningful database composed of 47 T2-weighted sagittal sequences from lumbar spine MRI. The patients were diagnosed with osteopenia or osteoporosis according to DEXA (patients with low bone mass). BEAUT achieved an accuracy of 92% and 97% AUC with feature selection to discriminate between patients from groups `Fractures' and `No Fractures'. The results support claiming that texture features potentially discriminate subjects with bone mass loss, spotting those at risk of fragility fractures.
Jonathan S. Ramos, Jamilly G. Maciel, Mirela Teixeira Cazzolato, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Agma J. M. Traina
CBMS3
2020 Semi-Automatic Ulcer Segmentation and Wound Area Measurement Supporting Telemedicine
abstract
Many patients suffer from chronic skin lesions, commonly known as ulcers. The size evolution of chronic wounds provides meaningful clues regarding the patient's clinical state for healthcare professionals and caretakers. Many studies have been proposed in recent years to support the treatment of skin ulcers. However, there is a lack of practical solutions, as existing studies are not targeted at immediate use in daily medical practice. In this work, we propose URule, an essentially practical framework for segmentation and measurement of skin ulcers. URule-App, a mobile instance of the framework, analyzes images taken by a common camera from a mobile device. The segmentation requires the user to manually outline the outsider region of both the wound and the measurement tool. URule-Seg segments the image and estimates the wound area. The user can further improve the estimated area by manually informing the span of a centimeter in the image. The experimental evaluation reveals that URule can accurately segment ulcer wounds semi-automatically, with an average F-Measure of 0.8 for segmentation, and processing measurement tools better than the manual process in three out of five tested rulers.
Mirela Teixeira Cazzolato, Jonathan S. Ramos, Lucas Santiago Rodrigues, Lucas C. Scabora, Daniel Y. T. Chino, Ana Elisa Serafim Jorge, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr., Agma J. M. Traina
CBMS1
2020 Taking Advantage of Highly-Correlated Attributes in Similarity Queries with Missing Values
Lucas Santiago Rodrigues, Mirela Teixeira Cazzolato, Agma J. M. Traina, Caetano Traina Jr.
SISAP2
2019 3DBGrowth: Volumetric Vertebrae Segmentation and Reconstruction in Magnetic Resonance Imaging
abstract
Segmentation of medical images is critical for making several processes of analysis and classification more reliable. With the growing number of people presenting back pain and related problems, the semi-automatic segmentation and 3D reconstruction of vertebral bodies became even more important to support decision making. A 3D reconstruction allows a fast and objective analysis of each vertebrae condition, which may play a major role in surgical planning and evaluation of suitable treatments. In this paper, we propose 3DBGrowth, which develops a 3D reconstruction over the efficient Balanced Growth method for 2D images. We also take advantage of the slope coefficient from the annotation time to reduce the total number of annotated slices, reducing the time spent on manual annotation. We show experimental results on a representative dataset with 17 MRI exams demonstrating that our approach significantly outperforms the competitors and, on average, only 37% of the total slices with vertebral body content must be annotated without losing performance/accuracy. Compared to the state-of-the-art methods, we have achieved a Dice Score gain of over 5% with comparable processing time. Moreover, 3DBGrowth works well with imprecise seed points, which reduces the time spent on manual annotation by the specialist.
Jonathan S. Ramos, Mirela Teixeira Cazzolato, Bruno S. Faiçal, Marcello Henrique Nogueira-Barbosa, Caetano Traina Jr., Agma J. M. Traina
CBMS2
2019 UCORM: Indexing Uncorrelated Metric Spaces for Concise Content-Based Retrieval of Medical Images
abstract
The large amount of medical exams generated by hospitals has a great potential to boost the support for physicians on decision making tasks. This requires efficient and reliable computational systems to retrieve relevant information in real-time. Existing Content-Based Image Retrieval (CBIR) systems rely on Metric Access Methods (MAMs) to speed-up the retrieval task. In this context, images are represented by Feature Extraction Methods (FEMs), according to information such as color or texture. However, MAMs usually index images based on a single FEM. Whenever physicians want to search for similar images using multiple FEMs simultaneously, they need to perform separated queries. In this work, we propose UCORM, an access method capable of indexing images using multiple FEMs by overlapping different metric spaces. UCORM selects the best FEMs to generate a concise yet accurate indexing space. It relies on an interesting use of Pearson correlation, that we named PCMS, to compute the correlation between different FEMs. PCMS allows UCORM to improve the retrieval task by minimizing the overlapping between metric spaces, resulting on fewer intermediary images when performing a query. Experimental analysis shows that UCORM prunes well the data distribution regions with low correlation between FEMs. Also, two medical application scenarios support our claim that UCORM is well-fitted for clinical environments.
Guilherme F. Zabot, Mirela Teixeira Cazzolato, Lucas C. Scabora, Bruno S. Faiçal, Agma J. M. Traina, Caetano Traina Jr.
CBMS2
2019 Efficient Indexing of Multiple Metric Spaces with Spectra
abstract
The widespread of social networks and online channels has increased the capture of large amounts of complex data, such as images and videos, which demand efficient and flexible tools to perform information retrieval. Many existing approaches to retrieve complex data follow the "Query by Similarity" paradigm, using Metric Access Methods (MAMs) to index complex data and speed-up information retrieval. In this context, many descriptors represent complex data using representative features such as color, shape, or texture for images. MAMs were initially designed to index features from complex data using only one descriptor, leading users to build several indexes when more than one descriptor is required. Recent approaches that use different representations in a single index structure suffer from a higher number of distance calculations. In this work, we propose the Spectra MAM, which indexes complex data using several features at once. Spectra integrates several metric spaces and answers queries based on one or more descriptors at once. Moreover, Spectra relies on existing correlations among different spaces to choose the best descriptors to obtain a concise yet accurate indexing space. Thus, it reduces the number of distance calculations, speeding up the query execution, and improving the resulting quality.
Guilherme F. Zabot, Mirela Teixeira Cazzolato, Lucas C. Scabora, Agma J. M. Traina, Caetano Traina Jr.
ISM2
2019 Employing Domain Indexes to Efficiently Query Medical Data From Multiple Repositories
abstract
Content-based retrieval still remains one of the main problems with respect to controversies and challenges in digital healthcare over big data. To properly address this problem, there is a need for efficient computational techniques, especially in scenarios involving queries across multiple data repositories. In such scenarios, the common computational approach searches the repositories separately and combines the results into one final response, which slows down the process altogether. In order to improve the performance of queries in that kind of scenario, we present the Domain Index, a new category of index structures intended to efficiently query a data domain across multiple repositories, regardless of the repository to which the data belong. To evaluate our method, we carried out experiments involving content-based queries, namely range and k nearest neighbor (kNN) queries, 1) over real-world data from a public data set of mammograms, as well as 2) over synthetic data to perform scalability evaluations. The results show that images from any repository are seamlessly retrieved, sustaining performance gains of up to 53% in range queries and up to 81% in kNN queries. Regarding scalability, our proposal scaled well as we increased 1) the cardinality of data (up to 59% of gain) and 2) the number of queried repositories (up to 71% of gain). Hence, our method enables significant performance improvements, and should be of most importance for medical data repository maintainers and for physicians' IT support.
Paulo H. Oliveira, Lucas C. Scabora, Mirela Teixeira Cazzolato, Willian D. Oliveira, Rafael S. Paixão, Agma J. M. Traina, Caetano Traina Jr.
IEEE J. Biomed. Health Informatics3
2018 ICARUS: Retrieving Skin Ulcer Images through Bag-of-Signatures
abstract
The images collected during medical exams are a strong asset for diagnosing and decision making. One scenario where clinical images are especially useful is the analysis of chronic lesions on the skin (skin ulcers). The visual appearance of these wounds may provide meaningful clues that may help physicians in the diagnosis. In this context, we propose ICARUS, an image retrieval system for dermatological ulcer images based on Bag-of-Visual-Words of color and texture signatures. ICARUS analyzes the image and extracts only the relevant signatures. The results show that ICARUS achieves improvement of up to 7% in image retrieval precision whereas being up to 5 orders of magnitude faster when compared to the state-of-the-art methods. Our results showed that ICARUS is effective and fast, and successfully adds semantic to the image representation.
Daniel Y. T. Chino, Lucas C. Scabora, Mirela Teixeira Cazzolato, Ana Elisa Serafim Jorge, Caetano Traina Jr., Agma J. M. Traina
CBMS3
2018 RAFIKI: Retrieval-Based Application for Imaging and Knowledge Investigation
abstract
Medical exams, such as CT scans and mammograms, are obtained and stored every day in hospitals all over the world, including images, patient data, and medical reports. It is paramount to have tools and systems to improve computer-aided diagnoses based on such huge volumes of stored information. The Content-Based Image Retrieval (CBIR) is a powerful paradigm to help reaching such a goal, providing physicians with intelligent retrieval tools to present him/her with similar or complementary cases, in which visual characteristics improve textual data. Employing comparative inspection on previous cases, the physician can obtain a more comprehensive understanding of the case he/she is working on. Current hospital systems do not carry native CBIR functionalities yet, relying on add-on subsystems, which often do not adhere to the existing relational database infrastructures. In this work, we propose RAFIKI, a software prototype that extends the Relational Database Management System (RDBMS) PostgreSQL, providing native support for CBIR functionalities, modular extensibility, and seamless integration for data science tools, such as Python and R. We show the applicability of our system by evaluating three clinical scenarios, performing queries over a real-world image dataset of lung exams. Our results spot actual potential in promoting informed decision-making from the physician's perspective. Besides, the system exhibited a higher performance when compared to previous systems found in the literature. Moreover, RAFIKI contributes with a model to establish how to put together CBIR concepts and relational data, providing a powerful design for further development of theoretical and practical concepts and tools.
Marcos Roberto Nesso Junior, Mirela Teixeira Cazzolato, Lucas C. Scabora, Paulo H. Oliveira, Gabriel Spadon, Jéssica Andressa de Souza, Willian D. Oliveira, Daniel Y. T. Chino, José F. Rodrigues Jr., Agma J. M. Traina, Caetano Traina Jr.
CBMS2
2018 Efficient and Reliable Estimation of Cell Positions
abstract
Sequences of microscopic images feature the dynamics of developing embryos. Automatically tracking the cells from such sequences of images allows understanding the dynamics which a living element demands to know its cells movement, which ideally should take place in real-time. The traditional tracking pipeline starts with image acquisition, data transfer, image segmentation to separate cells from the background, and then the actual tracking step. To speed up this pipeline, we hypothesize that a process capable of predicting the cell motion according to previous observations is useful. The solution must be accurate, fast and lightweight, and be able to iterate between the various components. In this work we propose CM-Predictor, which takes advantage of previous positions of cells to estimate their motion. When estimation takes place, we can omit costly acquisition, transfer and process of images, speeding up the tracking pipeline. The designed solution monitors the error of prediction, adapting the model whenever needed. For validation, we use four different datasets with sequences of images with developing embryos. Then we compare the estimated motion vectors of CM-Predictor with traditional tracking methods. Experimental results show that CM-Predictor is able to accurately estimate the motion vectors. In fact, CM-Predictor maintains the prediction quality of other algorithms and performs faster than them.
Mirela Teixeira Cazzolato, Agma J. M. Traina, Klemens Böhm
CIKM1
2017 BREATH: Heat Maps Assisting the Detection of Abnormal Lung Regions in CT Scans
abstract
Computed Tomography (CT) scans are often employed to diagnose lung diseases, as abnormal tissue regions may indicate whether proper treatment is required. However, detecting specific regions containing abnormalities in a CT scan demands time and effort of specialists. Moreover, different parts of a single lung image may present both normal and abnormal characteristics, what makes inaccurate the classification of a single lung as healthy (normal) or not. In this paper we propose the BREATH method, capable of detecting abnormalities in lung tissue regions, highlighting them by means of a heat map visualization. The method starts by segmenting lung tissues using a superpixel-based approach, followed by the training of a statistical model to represent normal tissues and, finally, the generation of a heat map showing abnormal regions that require attention from the physicians. We validated our statistical model using a dataset with 246 lung CT scans, where 40 are healthy and the remaining present varying diseases. Experimental results show that BREATH is accurate for lung segmentation with F-Measure of up to 0.99. The statistical modeling of healthy and abnormal lung regions has shown almost no overlap, and the detection of superpixels containing abnormalities presented precision values higher than 86%, for all values of recall. These values support our claim that the heat map representation of BREATH for the abnormal detection can be used as an intuitive method to assist physicians during the diagnosis.
Mirela Teixeira Cazzolato, Lucas C. Scabora, Alceu Ferraz Costa, Marcos Roberto Nesso Junior, Luis Fernando Milano Oliveira, Daniel S. Kaster, Caetano Traina Jr., Agma J. M. Traina
CBMS1
2017 Efficiently Indexing Multiple Repositories of Medical Image Databases
abstract
Performing content-based image retrieval over large repositories of medical images demands efficient computational techniques. The use of such techniques is intended to speed up the work of physicians, who often have to deal with information from multiple data repositories. When dealing with multiple data repositories, the common computational approach is to search each repository separately and merge the multiple results into one final response, which slows down the whole process. This can be improved if we build a mechanism able to search several repositories as if they were a single one, i.e. a mechanism to search the whole domain of medical images. Aiming at this goal, we propose the Domain Index, a new category of index structures aimed at efficiently searching domains of data, regardless of the repository to which they belong. To evaluate our proposal, we carried out experiments over multiple mammography repositories involving k Nearest Neighbor (kNN) and Range queries. The results show that images from any repository are seamlessly retrieved, even sustaining gains in performance of up to 36% in kNN queries and up to 7% in Range queries. The experimental evaluation shows that the Domain Index allows fast retrieval from multiple data repositories for medical systems, allowing a better performance in similarity queries over them.
Paulo H. Oliveira, Lucas C. Scabora, Mirela Teixeira Cazzolato, Willian D. Oliveira, Agma J. M. Traina, Caetano Traina Jr.
CBMS3
2017 Semantic Similarity Group By Operators for Metric Data
Natan A. Laverde, Mirela Teixeira Cazzolato, Agma J. M. Traina, Caetano Traina Jr.
SISAP2
2016 A Label-Scaled Similarity Measure for Content-Based Image Retrieval
abstract
Content-Based Image Retrieval (CBIR) has proven to be a suitable complement to traditional text-based searching. CBIR applications rely on two main steps, namely the representation of the images, and the similarity measuring between two represented images. Although modern segmentation and learning algorithms enable the accurate representation of local and global features within an image, how to properly compare the segmented objects is still an open issue. In this study, we propose a new comparison method called Counting-Labels Similarity Measure (CL-Measure). Our approach calculates the similarity between two images by comparing the labeled regions within these images and by balancing the influence of each label according to its predominance in both non-metric and metric fashion. The experiments on a real dataset of dermatological ulcers show that CL-Measure achieves a higher Precision for all values of Recall compared to its competitors in retrieval tasks.
Gustavo Blanco, Marcos V. N. Bedo, Mirela Teixeira Cazzolato, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina
ISM3
2013 A statistical decision tree algorithm for medical data stream mining
abstract
The use of computational resources can improve the diagnosis of medical diseases as a second opinion. Due to the large amount of data obtained daily, incremental techniques have been proposed to process medical data stream. In this paper we present an incremental decision tree classifier called StARMiner Tree (ST), which is based on Very Fast Decision Tree (VFDT) technique, to mine medical data. Different from VFDT, our proposed method ST does not depend on the number of reading samples to split a node. Because of it, ST is less conservative and describes the data since their first samples, being appropriate to be employed in medical environment, where not always a large number of data samples are available. We applied ST to four medical datasets, comparing the ST performance to the VFDT. The results indicated that ST is well-suited to deal with medical data streams, presenting high accuracy and low execution time.
Mirela Teixeira Cazzolato, Marcela X. Ribeiro
CBMS1