EDBT 2026 Demo / reviewers in the wild / expert
Marcos V. N. Bedo
dblp:133/1121 · also Marcos Bedo, Marcos Vinicius Naves Bedo
· DBLP profile ↗
24ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0003-2198-4670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable Reputation Management: A Multi-Task Prompting Approach Using Fine-Tuned PLMs for Sentiment and Topic ClassificationabstractSocial media platforms provide a direct, real-time channel for companies to engage with their audiences, making the management of corporate reputation on social media a pivotal factor for organizational success. Reputation is shaped by factors such as the relationship with the audience, crisis communication, public sentiment, and the topics most frequently associated with the company. Although reputation is often viewed as an intangible asset, it can be quantified and monitored through various metrics. Advances in AI and Pretrained Language Models (PLMs) have automated tasks such as sentiment analysis, sentiment strength assessment, and topic classification, creating new opportunities to manage reputation more efficiently. However, current PLM solutions typically address these tasks individually and require specialized training for each company, limiting scalability and flexibility. To address these challenges, we propose a novel system framework to assist public relations firms by leveraging PLMs for automatic classification. We evaluated this approach through experiments on four companies, comparing zero-shot and fine-tuned models in task-specific and multi-task configurations. Additionally, we explored the transfer of knowledge between client companies using fine-tuned models. Results indicate that multi-task, multi-company fine-tuned PLM models offer simpler system management with competitive performance compared to highly specialized models. Paulo R. S. da Costa Jr., Matheus Yasuo Ribeiro Utino, João Silva-Leite, Lyncoln S. de Oliveira, Rodrigo Salvador Monteiro, Rodrigo A. C. Dias, Daniel de Oliveira 0001, Paulo Mann, Marcos V. N. Bedo |
ICWSM | 9 |
| 2025 | PATSA-BIL: Pipeline for automated texture and structure analysis of borehole image logs
André M. Souza, Matheus A. Cruz, Paola M. C. Braga, Rodrigo B. Piva, Rodrigo A. C. Dias, Paulo R. Siqueira, Willian A. Trevizan, Candida M. de Jesus, Camilla Bazzarella, Rodrigo Salvador Monteiro, Flavia Bernardini, Leandro A. F. Fernandes, Elaine P. M. Sousa, Daniel de Oliveira 0001, Marcos V. N. Bedo |
Expert Syst. Appl. | 15 |
| 2024 | Aggregating embeddings from image and radiology reports for multimodal Chest-CT retrievalabstractThis paper proposes a multimodal retrieval system for Chest CT (called ChestFinder) that combines image and report embeddings’ into a filter-and-query strategy. ChestFinder is composed of three modules, namely (i) text transformation, (ii) feature extraction, and (iii) ranking aggregation. Text transformation is conducted by a fine-tuned Generative Pre-training Transformer model (GPT), and the image embeddings are extracted after (i) training a Residual Neural Network (ResNet-50), and (ii) reducing and scaling the encoded vectors. ChestFinder produces one list of similar images and another of related reports for each query input composed of a Chest CT with a radiology report. Then, the ChestFinder ranking aggregation module fuses those two lists to produce the final ordered set of retrieved objects, as in a top-k query. The aggregation is performed by a fine-tuned Threshold Algorithm (TA) whose weights are calculated by a Multi-Layer Perceptron (MLP) trained to label the reports. To examine the quality enhancement brought by this multimodal search, we constructed a dataset of Chest CTs from our University Hospital PACS/RIS systems by filtering distinct cases diagnosed with emphysema (one finding per case and with at least two radiologists agreeing on the diagnosis). A holdout experimental evaluation showed the ChestFinder search achieved higher Accuracy and Sensitivity than content-only top-k searches. Results also indicated quality gains drawn from the adjustments of ChestFinder modular components: (i) fine-tuned GPT achieved up to 0.89 F1-Score in data testing with a stable train/validation ratio for radiology reports, (ii) fine-tuned GPT significantly outperformed the zero-shot approach as well as a fine-tuned BERT, (iii) non-weighted ranking aggregation increased the search accuracy in up to 10%, and (iv) fine-tuned TA outperformed the baseline and non-weighted ranking aggregation in up to 52%. João Silva-Leite, Cristina A. P. Fontes, Alair S. Santos, Diogo G. Correa, Marcel Koenigkam-Santos, Paulo Mazzoncini de Azevedo Marques, Daniel de Oliveira 0001, Aline Paes, Marcos V. N. Bedo |
CBMS | 9 |
| 2024 | Enriching Hierarchical Navigable Small World Searches with Result Diversification
Mauro Weber, João Silva-Leite, Lúcio F. D. Santos, Daniel de Oliveira 0001, Marcos V. N. Bedo |
DEXA (1) | 5 |
| 2024 | A Computer Vision Model to Support Individuals with Disabilities Within University CampusesabstractThis study introduces MCOF, a multi-camera, computer vision-based system designed to assist visually impaired individuals with mobility on university campuses. The system operates locally and achieves the following performance metrics: (i) detecting body points within 1.2. 10–2seconds, (ii) identifying people, objects, and animals in 2.4. 10–2seconds, and (iii) detecting movements in 5.1. 10−2seconds. These results were obtained using a GTX 1,660 GPU with up to 6 cameras (or a 6,112MB stream) running concurrently. According to the MCOF architecture, events trigger tickets that are sent to an external information system, which can then implement its own safety and personnel protocols. Additionally, MCOF includes modules to handle electrical and network failures and features an obstruction detection routine for the cameras. Allan Costa Nascimento dos Santos, Karina de Paula, Marcos T. L. Vidal, João M. M. da Silva, Cledson Sousa, Leandro A. F. Fernandes, Tiago Bornia De Castro, Marcos V. N. Bedo, Troy C. Kohwalter, Carlos Alberto Malcher Bastos, Flávio Luiz Seixas, Natalia Castro Fernandes, Débora C. Muchaluat-Saade, George Ghinea |
HealthCom | 8 |
| 2023 | Adding Result Diversification to kNN-Based Joins in a Map-Reduce Framework
Vinícius Souza, Luiz Olmes Carvalho, Daniel de Oliveira 0001, Marcos V. N. Bedo, Lúcio F. D. Santos |
DEXA (1) | 4 |
| 2023 | Pushing diversity into higher dimensions: The LID effect on diversified similarity searching
Daniel L. Jasbick, Lúcio F. D. Santos, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Daniel de Oliveira 0001, Marcos V. N. Bedo |
Inf. Syst. | 6 |
| 2022 | Wia-Spine: A CBIR environment with embedded radiomic features to assess fragility fracturesabstractOsteoporosis is a systemic disorder that reduces the bone mineral density, increasing the vertebrae's fragility and proneness to fracture. Although the bone densitometry index t-Score is a solid marker for the osteoporosis diagnosis, its measure alone is insufficient to predict the future development of fragility fractures. A complementary approach to address vertebral bone characterization is the analysis of magnetic resonance imaging (MRI) by radiomic features, which model vertebral bodies' morphological properties after color and texture. Radiomic features have been employed for detecting fragility fractures in related work, but, to the best of our knowledge, no study has been conducted on their suitability to recover similar, diagnosed cases that could hint at future fractures. We fulfill this gap by designing a Content-based Image Retrieval (CBIR) tool with embedded radiomic features, which uses past cases recovered from an annotated database to (i) identify an existing fragility fracture in a query vertebra and (ii) predict a fracture to a query vertebra from an aging patient. The proposed CBIR was evaluated on a reference database of 273 vertebral bodies from sagittal T2-weighted MRIs. The results indicate our fine-tuned approach spotted fragility fractures accurately$(\mathrm{F}1-\text{Score} =0.83,\ \text{Precision} =0.83,\ \text{AUC} =0.81,\ \text{CI} =95\%)$. We also investigated the CBIR potential to predict fractures in a case study regarding three patients from the reference database (confirmed osteoporosis, MRI in [2012–2017]). The system correctly inferred the prediction of future fractures for query vertebrae, which were confirmed a few years later (MRI in [2018–2021]). Such empirical findings suggest CBIR can support a differential diagnosis in the assessment of local fragility fractures. Marcos V. N. Bedo, Jonathan S. Ramos, Agma J. M. Traina, Caetano Traina Jr., Marcello Henrique Nogueira-Barbosa, Paulo Mazzoncini de Azevedo Marques |
CBMS | 1 |
| 2020 | Some Branches May Bear Rotten Fruits: Diversity Browsing VP-Trees
Daniel L. Jasbick, Lúcio F. D. Santos, Daniel de Oliveira 0001, Marcos V. N. Bedo |
SISAP | 4 |
| 2020 | Capturing and Analyzing Provenance from Spark-based Scientific Workflows with SAMbA-RaP
Thaylon Guedes, Lucas Bertelli Martins, Maria Luiza Furtuozo Falci, Vítor Silva 0003, Kary A. C. S. Ocaña, Marta Mattoso, Marcos V. N. Bedo, Daniel de Oliveira 0001 |
Future Gener. Comput. Syst. | 7 |
| 2019 | A Two-Phase Learning Approach for the Segmentation of Dermatological WoundsabstractTissue segmentation in photographs of lower limb chronic ulcers is a non-intrusive approach that supports dermatological analyses. This paper presents 2PLA, a method that combines supervised and unsupervised learning strategies for enhancing the segmentation of dermatological wounds. Given an ulcer photo captured according to a fixed protocol, 2PLA first phase performs a pixelwise classification of points of interest, whereas pre-processing filters are employed for the smoothing of image noise. The cleaned image is further sent to the 2PLA divide-and-conquer second phase. It builds upon SLIC superpixel construction algorithm for dividing the lower limb into regions of interest with well-defined borders, and clusters the superpixels by taking advantage of the similarity-based DBSCAN algorithm. We set up the phases of our method by using a real annotated set of dermatological wounds, and empirical evaluations on representative samples up to 100,000 points showed a compact Multi-Layer Perceptron with Levenberg-Marquardt training algorithm (Cohen-Kappa = .971, Sensitivity = .98, and Specificity = .98) outperformed other classifiers as 2PLA first phase. Additionally, experimental trials on DBSCAN with five distance functions (L1, L2, L∞, Canberra, and BrayCurtis) indicated L1function provided fewer groups in comparison to the competitors, and the number of clusters was an exponential decay to the similarity ratio. Accordingly, we used the elbow criterion for finding the L1-based DBSCAN threshold as 2PLA second phase parameterization. We evaluated the fine-tuned setting of our method over a labeled set of ulcer images, and wounded tissues were segmented within a .05 Mean Absolute Error ratio. These results illustrate the impact of learning parameters on 2PLA as well as the method efficacy for wound segmentation. Wellington S. Silva, Daniel L. Jasbick, Rodrigo Erthal Wilson, Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Daniel de Oliveira 0001, Marcos V. N. Bedo |
CBMS | 9 |
| 2019 | A k-Skyband Approach for Feature Selection
Marcos V. N. Bedo, Paolo Ciaccia, Davide Martinenghi, Daniel de Oliveira 0001 |
SISAP | 1 |
| 2018 | Exploring Diversified Similarity with KundahaabstractExploring large medical image sets by means of traditional similarity query criteria (e.g., neighborhood) can be fruitless if retrieved images are too similar among themselves. This demonstration introduces Kundaha, an exploration tool that assists experts in retrieving and navigating on results from a diversified similarity perspective of user-posed queries. Its implementation includes a wide set of metrics, descriptors, and indexes for enhancing query execution. Users can combine such features with diversified similarity criteria for the organized exploration of result sets and also employ relevance feedback cycles for finding new query-based viewpoints. Lúcio F. D. Santos, Gustavo Blanco, Daniel de Oliveira 0001, Agma J. M. Traina, Caetano Traina Jr., Marcos V. N. Bedo |
CIKM | 6 |
| 2018 | Standard SQL Approaches for Similarity SearchingabstractThis paper addresses complex data storage and retrieval in RDBMS, which depends on metric distance functions for the assessment of data dissimilarity. However, both the empirical analysis of strategies for complex data storage and the definition of a suitable representation for similarity query operators are still open issues in the literature. Here, we fulfill those gaps through the classification, implementation, and evaluation of existing approaches for complex data storage according to four structures found in standard SQL, namely relational, object-relational, binary and semi-structured. Moreover, we also discuss a comprehensive model for complex data retrieval, whose conception of similarity operators is consistent with standard SQL representations. Accordingly, a distance function representation is presented, which enables the RDBMS query processor to interpret and execute physical similarity operators. Experimental results indicate: (i) relational and object-relational structures outperform the other two competitors in the majority of scenarios, whereas (ii) object-relational strategy enables the use of a broader representation. Pedro Henrique Braga Siqueira, Paulo H. Oliveira, Marcos V. N. Bedo, Daniel S. Kaster |
CLEI | 3 |
| 2018 | What Lies Beyond Structured Data? A Comparison Study for Metric Data Storage
Pedro Henrique Braga Siqueira, Paulo H. Oliveira, Marcos V. N. Bedo, Daniel S. Kaster |
DEXA (2) | 3 |
| 2018 | The Merkurion approach for similarity searching optimization in Database Management SystemsabstractModern Database Management Systems (DBMSs) retrieve songs that resemble those in a music dataset, identify plagiarism in a set of documents, or provide past cases to physicians by taking into account the characteristics of a query exam. All such tasks require the comparison of data by similarity, which can be expressed in terms of distance-based queries in metric spaces. Traditional query processing relies mostly on histograms for describing the data distribution space and choosing a data retrieval path that quickly leads to the answer, discarding comparisons of most unwanted data. However, DBMSs still lack adequate support for selectivity estimation of query operators for data types embedded in metric spaces. This article addresses a novel strategy that extends the query optimizer of a DBMS, so that it can also perform both logical and physical query plan optimizations in searches that include similarity predicates. The proposal, named Merkurion, updates the concept of Data Distribution Space and captures data distributions according to the distances between the elements within a dataset. Moreover, it employs concise representations of such distributions, called synopses, for the definition of rules that enable similarity searching optimization. An extensive evaluation of Merkurion in real-world datasets has proven its effectiveness and broad applicability to many data domains. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
Data Knowl. Eng. | 1 |
| 2016 | A Label-Scaled Similarity Measure for Content-Based Image RetrievalabstractContent-Based Image Retrieval (CBIR) has proven to be a suitable complement to traditional text-based searching. CBIR applications rely on two main steps, namely the representation of the images, and the similarity measuring between two represented images. Although modern segmentation and learning algorithms enable the accurate representation of local and global features within an image, how to properly compare the segmented objects is still an open issue. In this study, we propose a new comparison method called Counting-Labels Similarity Measure (CL-Measure). Our approach calculates the similarity between two images by comparing the labeled regions within these images and by balancing the influence of each label according to its predominance in both non-metric and metric fashion. The experiments on a real dataset of dermatological ulcers show that CL-Measure achieves a higher Precision for all values of Recall compared to its competitors in retrieval tasks. Gustavo Blanco, Marcos V. N. Bedo, Mirela Teixeira Cazzolato, Lúcio F. D. Santos, Ana Elisa Serafim Jorge, Caetano Traina Jr., Paulo Mazzoncini de Azevedo Marques, Agma J. M. Traina |
ISM | 2 |
| 2016 | When Similarity is Not Enough, Ask for Diversity: Grouping Elements Based on InfluenceabstractCrowdsourcing images have been increasingly employed for mapping emergency scenarios, which helps rescue forces in choosing contingency plans. In this scenario, similarity searching can be used to retrieve related images from past situations. However, the retrieved images often are similar among themselves and, therefore, add little to none new information to the rescue decision-making process. In this paper, we take advantage of diversity queries to increase the variety of the representative elements about an incident, whereas the remaining and related data are grouped according to the set of representatives. Thus, our approach enables content retrieval, grouping and an easier exploration of the result set. Experiments performed on real datasets shows that our proposal outperforms the existing methods regarding both quality and performance, being at least three orders of magnitude faster. Lúcio F. D. Santos, Luiz Olmes Carvalho, Marcos V. N. Bedo, Agma J. M. Traina, Caetano Traina Jr. |
ISM | 3 |
| 2015 | Color and Texture Influence on Computer-Aided Diagnosis of Dermatological UlcersabstractThis study presents an analysis of classification techniques for Computer-Aided Diagnosis (CAD) regarding ulcerated lesions. We focus on determining influence of both color and texture in the automated image classification and its implication. To do so, we assayed a dataset of dermatological ulcers containing five variations in terms of tissue composition of lesion skin: granulation (red), fibrin (yellow), callous (white), necrotic (black), and a mix of the previous variations (mixed). Every image was previously labelled by experts regarding this red-yellow-black-white-mixed model. We employed specially designed color and texture extractors to represent the dataset images, namely: Color Layout, Color Structure, Scalable Color, Edge Histogram, Haralick, and Texture-Spectrum. The first three are color feature extractors and the last three are texture extractors. Following, we employed the Symmetrica Uncert Attribute Eval method to determine the features suitable for image classification. We tested a set of classifiers that follows distinct paradigms over the selected features, achieving an accuracy ratio of up to 77% in terms of images correctly classified, with the area under the receiver operating characteristic (ROC) curve up to 0.84. The classification performance and the selected features enabled us to determine that texture features were more predominant than color in the entire classification process. Marcos V. N. Bedo, Lúcio F. D. Santos, Willian D. Oliveira, Gustavo Blanco, Agma J. M. Traina, Marco Antonio Frade, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr. |
CBMS | 1 |
| 2015 | Compact distance histogram: a novel structure to boost k-nearest neighbor queriesabstractThe k-Nearest Neighbor query (k-NNq) is one of the most useful similarity queries. Elaborated k-NNq algorithms depend on an initial radius to prune regions of the search space that cannot contribute to the answer. Therefore, estimating a suitable starting radius is of major importance to accelerate k-NNq execution. This paper presents a new technique to estimate a tight initial radius. Our approach, named CDH-kNN, relies on Compact Distance Histograms (CDHs), which are pivot-based histograms defined as piecewise linear functions. Such structures approximate the distance distribution and are compressed according to a given constraint, which can be a desired number of buckets and/or a maximum allowed error. The covering radius of a k-NNq is estimated based on the relationship between the query element and the CDHs' joint frequencies. The paper presents a complete specification of CDH-kNN, including CDH's construction and radii estimation. Extensive experiments on both real and synthetic datasets highlighted the efficiency of our approach, showing that it was up to 72% faster than existing algorithms, outperforming every competitor in all the setups evaluated. In fact, the experiments showed that our proposal was just 20% slower than the theoretical lower bound. Marcos V. N. Bedo, Daniel S. Kaster, Agma J. M. Traina, Caetano Traina Jr. |
SSDBM | 1 |
| 2014 | MSSF: A Step towards User-Friendly Multi-cloud Data DispersalabstractWith an increasing number of companies and individuals adopting cloud computing for their data needs. Naturally, there is a shift in financial and operational costs and and provided services should meet users' performance and cost expectations. We focus on storage and propose MSSF, a Multi-cloud Storage Selection Framework. MSSF contains a basic set of algorithms, a set of security rules and a formal definition of user profiles allowing to fit cloud storage services to user needs. Preliminary experiments investigate the cost differences between two baseline algorithms and user profile models. Considering the promising initial results we provide several observations about the MSSF decision making process that will help with future improvements. Rafael Mira De Oliveira Libardi, Marcos V. N. Bedo, Stephan Reiff-Marganiec, Júlio Cezar Estrella |
IEEE CLOUD | 2 |
| 2014 | Being Similar is Not Enough: How to Bridge Usability Gap through Diversity in Medical ImagesabstractIn this paper we present a technique developed to bridge the usability gap in Content-Based Medical Image Retrieval (CBMIR) systems exploring both similarity and diversity. Usability gaps are related to how easy to use a software tool from the radiologist's perspective is. Although much have been done to better express similarity queries, the use of CBMIR over massive databases may have drawbacks that impact its usability. We claim that much of the problems derives from the fact that many images returned are closer to each other than to the query element (near-duplicates). To target this nuisance, we propose to boost similarity queries with diversity, using a technique to hierarchically cluster near-duplicates. We tailored a domain-independent and parameter-free method by controlling the maximum area reached in the search space. This novel approach to improve CBMIR systems take advantage of diversity expectations. The proposed approach BridGE (Better result with influence diversification to Group Elements) aims at adding new relevant information to the analysts, reducing the need of further query refinement or relevance feedback cycles. The results are displayed to the specialist as a traditional CBMIR result whereas the radiologists are able to expand the clusters and navigate through them. The results support our claim that a CBMIR system empowered with diversity is able to bridge the usability gap, grouping near-duplicates and being at least 2 orders of magnitude faster than its mainly competitors. Lúcio F. D. Santos, Marcos V. N. Bedo, Marcelo Ponciano-Silva, Agma J. M. Traina, Caetano Traina Jr. |
CBMS | 2 |
| 2013 | Does a CBIR system really impact decisions of physicians in a clinical environment?abstractContent-based image retrieval systems are employed in several areas. One of the most prominent area is the medical field, due to the huge volume of digital images daily generated in healthcare institutions employed for decision making. There are several works applying CBIR techniques over medical images. However, the great majority of them do not verify whether the systems are actually considered by the specialists as a pontential aid in a real environment. In order to fill this research void in the literature, this work explores user experiments in a CBIR system involving resident physicians and radiologists. To do so, we developed a CBIR system according to requirements provided by the specialists and employed a methodology to analyze the effectiveness of the system for supporting them in clinical routine. The methodology aims at evaluating the system's impact in the user's decision, inquiring the specialists about the image classification and their degree of certainty in different situations using the system. By analyzing the obtained results we can argue that the proposed methodology joined with our medical CBIR system presented a high acceptance and viability rate regarding the radiologists interests in the clinical practice domain, providing a novel approach to analyze CBIR systems under realistic conditions. Marcelo Ponciano-Silva, Juliana P. Souza, Pedro Henrique Bugatti, Marcos V. N. Bedo, Daniel S. Kaster, Rosana T. V. Braga, Angela D. Bellucci, Paulo Mazzoncini de Azevedo Marques, Caetano Traina Jr., Agma J. M. Traina |
CBMS | 4 |
| 2013 | A Similarity-Based Approach for Financial Time Series Analysis and Forecasting
Marcos V. N. Bedo, Davi Pereira dos Santos, Daniel S. Kaster, Caetano Traina Jr. |
DEXA (2) | 1 |