VLDB 2026 Research / reviewers in the wild / expert
Jorge Miguel 0002
dblp:121/2117-2 · also Jorge Miguel Ferreira da Silva, Jorge Miguel Silva
· DBLP profile ↗
16ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-6331-6091ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 5 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Biochef: a client-side WebAssembly-based workflow builder for genomic data analysisabstractAbstract Background Genomics analyses often rely on command-line tools executed via remote servers, imposing usability barriers for non-technical users and raising privacy concerns. WebAssembly (WASM) enables native-code execution directly in web browsers, eliminating installations and data transfers. Results We introduce BioChef, a client-side genomic workflow platform that uses WASM. BioChef compiles a genomics toolkit into browser-executable modules and exposes them through a drag-and-drop GUI designed to be intuitive. The system provides real-time validation, flexible input methods (form-based and JSON), intermediate step inspections, and reproducible workflows exportable as bash scripts or configuration files. Performance benchmarks across major browsers (Chromium, Gecko, WebKit) demonstrate rapid initialization (LCP 0.583 s), responsive interactivity (INP 30.5 ms), minimal layout shifts (CLS 0.01), and acceptable overhead (average 181.5 ms initial WASM module load). Although browser execution introduced performance penalties ( $$\sim $$ ∼ 130 $$\times $$ × slower than native), BioChef workflows still significantly outperformed traditional web services such as Galaxy by avoiding network delays and server-side queueing (11.3 $$\times $$ × faster in a standard pipeline benchmark). Conclusions BioChef shows how WebAssembly on the client side can democratize genomic data processing, ensuring privacy, reproducibility and ease of use without external dependencies. To our knowledge, this is the first fully client-side, graphical genomic workflow environment powered by WASM. Joaquim Vertentes Rosa, João Andrade, Jorge Miguel 0002, José Luís Oliveira |
BMC Bioinform. | 3 |
| 2025 | A Federated Random Forest Solution for Secure Distributed Machine LearningabstractPrivacy and regulatory barriers often hinder centralized machine learning solutions, particularly in sectors like healthcare where data cannot be freely shared. Federated learning has emerged as a powerful paradigm to address these concerns; however, existing frameworks primarily support gradientbased models, leaving a gap for more interpretable, tree-based approaches. This paper introduces a federated learning framework for Random Forest classifiers that preserves data privacy and provides robust performance in distributed settings. By leveraging PySyft for secure, privacy-aware computation, our method enables multiple institutions to collaboratively train Random Forest models on locally stored data without exposing sensitive information. The framework supports weighted model averaging to account for varying data distributions, incremental learning to progressively refine models, and local evaluation to assess performance across heterogeneous datasets. Experiments on two realworld healthcare benchmarks demonstrate that the federated approach maintains competitive predictive accuracy—within a maximum 9% margin of centralized methods—while satisfying stringent privacy requirements. These findings underscore the viability of tree-based federated learning for scenarios where data cannot be centralized due to regulatory, competitive, or technical constraints. The proposed solution addresses a notable gap in existing federated learning libraries, offering an adaptable tool for secure distributed machine learning tasks that demand both transparency and reliable performance. The tool is available at https://github.com/ieeta-pt/fed_rf Alexandre Cotorobai, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 2 |
| 2024 | HealthDBFinder: a question-answering task for health database discoveryabstractIntegrating advanced data processing technologies into healthcare has shifted the medical studies paradigm. These evolve from data collection into management and analysis of Electronic Health Records (EHR) data. This change improved patient care and expanded the scope of clinical research through the secondary usage of existing data. Even though this problem was already solved in other initiatives, it raised new challenges, namely regarding cohort definition, data discovery, and evaluating the study feasibility. There are database catalogues to help in those tasks, but these fail in some cases due to insufficient information. Therefore, in this paper, we address this challenge by proposing a baseline method for information retrieval, including a synthetic dataset for further research. The information present in the dataset was generated from metadata extracted from real-world databases, which represents real problems that do not yet have a solution. The source code of this work is available at http://github.com/bioinformatics-ua/HealthDBFinder. João Rafael Almeida, Jorge Miguel 0002, Luís Carlos Afonso, Tiago Melo Almeida, Rui Antunes 0002, Richard Adolph Aires Jonker, João António Reis, Dimitri Alexandre da Silva, Sérgio Matos, José Luís Oliveira |
CBMS | 2 |
| 2024 | A comprehensive study of databases to assess the reliability of metagenomic toolsabstractThe advancement of metagenomics is closely tied to bioinformatic tools. These tools, which are essential for taxonomic classification and functional annotation, derive their reliability from the databases they use. However, the challenge arises when comparing these tools through literature, as their evaluations often occur within specific datasets. This practice makes it difficult to compare them directly with state-of-the-art tools, as it obscures how they might perform across a broader range of data. In this study, we evaluate the suitability of different databases for assessing the viability of metagenomic tools. By assessing these databases and comparing them to general databases, we aim to provide researchers with valuable insights into selecting the most appropriate database for their metagenomic studies, enhance the reliability and reproducibility of metagenomic technologies, and overcome challenges such as resolving low-abundance species, distinguishing closely related species, or handling environmental samples with high diversity. Inês Branco Martins, Jorge Miguel 0002, João Rafael Almeida |
CIBCB | 2 |
| 2024 | Ahead of Time Prediction of Decorated Particleboard Production Disruptions and Defects Using Single and Multi-Target AutoMLabstractThis paper proposes a Machine Learning (ML) approach to perform an Ahead-of-Time (AoT) prediction of decorated particle-board production disruptions and defects. We worked with a Portuguese company that is adopting the Industry 4.0 concept aiming to improve their decorated particleboard production planning (e.g., reducing production time and waste of materials). This company’s business needs are addressed in terms of two nontrivial binary Classification tasks (production disruptions and defects). The AoT prediction is achieved by using only input attributes available before the execution of the production process. To reduce the modeling effort, we focus on Automated ML (AutoML) methods, under two main approaches: Single-Target Classification (STC) and Multi-Target Classification (MTC). The former is achieved by adopting the popular H2O AutoML tool, while the latter adopts a deep learning neural network automatically tuned by using a Bayesian search. The computational experiments adopted a realistic rolling window evaluation over recently collected industrial data (comprising 14 months). Overall, interesting predictive results were achieved by both AutoML approaches, outperforming a baseline Decision Tree method. In addition, an eXplainable Artificial Intelligence (XAI) method based on a Sensitivity Analysis (SA) was adopted, allowing the identification of the most relevant inputs, which is valuable knowledge to support the decorated particleboard production planning. Arthur Matta, Luís Miguel Matos, André Luiz Pilastri, Jorge Miguel 0002, Miguel Bastos Gomes, Paulo Cortez 0001 |
KES | 4 |
| 2024 | Enhancing metagenomic classification with compression-based featuresabstractMetagenomics is a rapidly expanding field that uses next-generation sequencing technology to analyze the genetic makeup of environmental samples. However, accurately identifying the organisms in a metagenomic sample can be complex, and traditional reference-based methods may need to be more effective in some instances. In this study, we present a novel approach for metagenomic identification, using data compressors as a feature for taxonomic classification. By evaluating a comprehensive set of compressors, including both general-purpose and genomic-specific, we demonstrate the effectiveness of this method in accurately identifying organisms in metagenomic samples. The results indicate that using features from multiple compressors can help identify taxonomy. An overall accuracy of 95% was achieved using this method using an imbalanced dataset with classes with limited samples. The study also showed that the correlation between compression and classification is insignificant, highlighting the need for a multi-faceted approach to metagenomic identification. This approach offers a significant advancement in the field of metagenomics, providing a reference-less method for taxonomic identification that is both effective and efficient while revealing insights into the statistical and algorithmic nature of genomic data. The code to validate this study is publicly available at https://github.com/ieeta-pt/xgTaxonomy. Jorge Miguel 0002, João Rafael Almeida |
Artif. Intell. Medicine | 1 |
| 2023 | A FAIR Approach to Real-World Health Data Management and AnalysisabstractThe increasing of health data sources to support clinical practice is opening the path for its secondary use in biomedical research. This changes the research paradigm, from data generation to data management and analysis. Although the potential for secondary use of this data is vast, including the improvement of healthcare systems and the advancement of clinical research, data discovery is challenging. In order to maximize data reusability, the FAIR principles have been developed as a guiding framework for system development. Nevertheless, the discovery and reuse of biomedical data present two main challenges: i) data partners grappling with ethical and social concerns related to data discoverability; ii) clinical researchers struggling to find the best data sources for their research studies. In this paper, we present a platform that provides a set of tools, compliant with the FAIR principles, to help data custodians when sharing data about biomedical databases, while allowing researchers to search for and select databases that meet their specific research needs. João Rafael Almeida, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 2 |
| 2023 | SecureFASTA: Ensuring privacy and trust when sharing genomic dataabstractGenomics has profoundly influenced the field of medicine, with advancements in DNA sequencing contributing to personalized medicine and a more comprehensive understanding of various diseases' genomic underpinnings. Sharing genomic data is vital for progressing the field and devising novel approaches to decipher the genome. Nevertheless, the sensitive nature of this information necessitates robust security measures for protection during storage and transfer. In this paper, we introduce SecureFASTA, a novel tool for securely encrypting and decrypting FASTA files without requiring a shared secret while minimizing the number of keys exchanged between pairs. Our approach combines symmetric and asymmetric encryption techniques, utilizing the Advanced Encryption Standard (AES) cypher and Rivest-Shamir-Adleman (RSA) encryption. Additionally, we implement a checksum function using the Secure Hash Algorithm (SHA-256) to verify the integrity of transferred FASTA files. Our evaluation demonstrates that SecureFASTA is fast, reliable, and secure, surpassing existing tools in terms of security and user-friendliness. Consequently, it offers a valuable solution for securely sharing and leveraging sensitive genomic data, marking a significant breakthrough in genomics. The tool's source code is available at https://github.com/bioinformatics-ua/SecureFASTA. Diniz Cruz, João Rafael Almeida, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 3 |
| 2022 | The value of compression for taxonomic identificationabstractAdvances in DNA sequencing technologies have led to an unprecedented growth of sequenced data. However, when sequencing de-novo genomes, one of the biggest challenges is the classification of DNA sequences that do not match with any biological sequence from the literature. The use of reference-free methods to identify these organisms supported by compressors is one strategy for taxonomic identification. However, with the high number of compressors available, and the computational resources required to operate them, there is a problem in selecting the best compressors for classification with limited computational resources. In this paper, we present a two-step pipeline to analyze nine compressors, to understand which ones could be the best candidates for taxonomic identification. We use 500 randomly selected sequences from five taxonomic groups to conduct this analysis. The results show that besides being an excellent repre-sentative feature, depending on the compressor, the Normalized Compression (NC) reflects different aspects concerning the nature of a given sequence and its complexity. Furthermore, we show that neither the compression capability of a compressor nor the compressibility of the file correlates with classification accuracy. The code used in this work is publicly available at https://github.com/bioinformatics-ua/COMPACT. Jorge Miguel 0002, João Rafael Almeida |
CBMS | 1 |
| 2021 | Persistent minimal sequences of SARS-CoV-2abstractMOTIVATION: Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. RESULTS: Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. AVAILABILITY AND IMPLEMENTATION: https://github.com/cobilab/eagle. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Diogo Pratas, Jorge Miguel 0002 |
Bioinform. | 2 |
| 2021 | Automatic analysis of artistic paintings using information-based measures
Jorge Miguel 0002, Diogo Pratas, Rui Antunes 0002, Sérgio Matos, Armando J. Pinho |
Pattern Recognit. | 1 |
| 2020 | Dicoogle Framework for Medical Imaging Teaching and ResearchabstractOne of the most noticeable trends in healthcare over the last years is the continuous growth of data volume produced and its heterogeneity. In the medical imaging field, the evolution of digital systems is supported by the PACS concept and the DICOM standard. These technologies are deeply grounded in medical laboratories, supporting the production and providing healthcare practitioners with the ability to set up collaborative work environments with researchers and academia to study and improve healthcare practice. However, the complexity of those systems and protocols makes difficult and time-consuming to prototype new ideas or develop applied research, even for skilled users with training in those environments. Dicoogle emerges as a reference tool to achieve those objectives through a set of resources aggregated in the form of a learning pack. It is an open-source PACS archive that, on the one hand, provides a comprehensive view of the PACS and DICOM technologies and, on the other hand, provides the user with tools to easily expand its core functionalities. This paper describes the Dicoogle framework, with particular emphasis in its Learning Pack package, the resources available and the impact of the platform in research and academia. It starts by presenting an overview of its architectural concept, the most recent research backed up by Dicoogle, some remarks obtained from its use in teaching, and worldwide usage statistics of the software. Moreover, a comparison between the Dicoogle platform and the most popular open-source PACS in the market is presented. Rui Lebre, Eduardo Pinho, Jorge Miguel 0002, Carlos Costa 0001 |
ISCC | 3 |
| 2018 | Face De-Identification Service for Neuroimaging VolumesabstractDigital medical imaging is a fundamental tool for improving medical practice workflows and supporting clinical diagnosis. Nowadays, healthcare institutions are usually very supported by information and communication systems that meet regular practice requirements. However, the usage of those platforms in collaborative, research and educational scenarios faces several problems. One of the key issues is related with patient data privacy, namely with concerns related with the visual anonymization of studies. In the neuroimaging field, this subject is more complex since, even after removing the patient's information from the images meta-data or burned in the pixel data, it is still possible to identify the patients through 3D reconstruction of the volume. This article proposes and describes the implementation of an end-user service that allows neuroimages facial de-identification of CT volumes, being fully interoperable with production repositories. The solution was validated using a public dataset and made available to the community through its integration with an open source archive server. Jorge Miguel 0002, António Guerra, João Figueira Silva, Eduardo Pinho, Carlos Costa 0001 |
CBMS | 1 |
| 2018 | Ejection Fraction Classification in Transthoracic Echocardiography Using a Deep Learning ApproachabstractCardiovascular diseases are the leading cause of death worldwide. These diseases are related with a broad range of factors but usually show high correlation with diminished left ventricle function, which can be evaluated by measuring the ventricular ejection fraction through transthoracic echocardiography (TTE), a cost-effective and highly portable first-line diagnosing technique. Ejection fraction (EF) is currently determined through a semi-automatic process that requires manual delineation of the left ventricle area both in a diastolic and systolic frame of the patient's exam. To remove this manual annotation step, which is both time-consuming and user dependent, automatic Computer-Aided Diagnosis (CAD) systems can be used. Herein, we propose the first steps for such a system that classifies ejection fraction in four classes, based on TTE exams, with the objective of automatically providing valuable information to physicians. Our classification method is based on a 3D-Convolutional Neural Network (3D-CNN) trained on a dataset constructed with exams from a cardiology reference center. The dataset creation consisted of three main steps: firstly, for each exam, cine-loops showing the apical 4 chambers view were manually selected; then, 30 sequential frames were extracted from each cine-loop; finally, each frame was pre-processed to mask burned-in metadata. The neural network was designed to explore concepts such as convolutions using asymmetric filters and residual learning blocks. The model was trained on a dataset with 4000 TTE exams and tested on a separate dataset containing 1600 TTE cases. We obtained an accuracy of 78% and a F1 score of 71.3% for unhealthy EF (below 45%), 63.3% for intermediate EF (45-55%), 72.3% for healthy EF (55-75%) and 54.6% for abnormally high EF (above 75%). These results are promising and show that convolutional neural networks can be applied to this domain. Furthermore, this work will serve as a foundation for future research where other relevant cardiac metrics will be determined. João Figueira Silva, Jorge Miguel 0002, António Guerra, Sérgio Matos, Carlos Costa 0001 |
CBMS | 2 |
| 2018 | Reactive Through Services - Opinionated Framework for Developing Reactive Services
Micael Pedrosa, Jorge Miguel 0002, Carlos Costa 0001 |
CLOSER | 2 |
| 2018 | Controlled searching in reversibly de-identified medical imaging archives
Jorge Miguel 0002, Eduardo Pinho, Eriksson J. Melicio Monteiro, João Figueira Silva, Carlos Costa 0001 |
J. Biomed. Informatics | 1 |