EDBT 2026 Demo / reviewers in the wild / expert
José Luís Oliveira
dblp:10/1144
· DBLP profile ↗
75ranked-venue papers
1as first author
32since 2021 · last 2026
0000-0002-6672-6176ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 56 · 22 since 2021Artificial intelligence and machine learning · 32 · 21 since 2021Human-computer interaction and ubiquitous computing · 28 · 17 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Computer networks · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Data Management for Smart Campus Digital Twin Applications
Luís Carlos Afonso, João Rafael Almeida, José Luís Oliveira |
DATA (1) | 3 |
| 2026 | A Serverless Client-Side Privacy Index for Sensitive Data Processing
José Gameiro, José Luís Oliveira, João Rafael Almeida |
DATA (2) | 2 |
| 2026 | Optimizing Job Rotation in Assembly Lines by Balancing Productivity and Worker Well-Being
Joana Rafaela Almeida, João Rafael Almeida, José Luís Oliveira |
ICORES | 3 |
| 2026 | Managing Cybersecurity Compliance with Structured Guidance and Integrated Audit Support
Mariana Andrade, João Rafael Almeida, José Luís Oliveira |
SECRYPT (2) | 3 |
| 2026 | A Predictive Data-Driven Framework for Multi-Line Manufacturing Throughput Analysis
Joana Rafaela Almeida, Raquel Paradinha, Luís Carlos Afonso, João Rafael Almeida, José Luís Oliveira |
SIMULTECH | 5 |
| 2026 | A Governance and Architectural Framework for Agentic AI-Driven Gamified Security Awareness
Luís Filipe Gomes, Luís Miguel Batista, António Deus, José Luís Oliveira, João Rafael Almeida |
SIMULTECH | 4 |
| 2026 | Biochef: a client-side WebAssembly-based workflow builder for genomic data analysisabstractAbstract Background Genomics analyses often rely on command-line tools executed via remote servers, imposing usability barriers for non-technical users and raising privacy concerns. WebAssembly (WASM) enables native-code execution directly in web browsers, eliminating installations and data transfers. Results We introduce BioChef, a client-side genomic workflow platform that uses WASM. BioChef compiles a genomics toolkit into browser-executable modules and exposes them through a drag-and-drop GUI designed to be intuitive. The system provides real-time validation, flexible input methods (form-based and JSON), intermediate step inspections, and reproducible workflows exportable as bash scripts or configuration files. Performance benchmarks across major browsers (Chromium, Gecko, WebKit) demonstrate rapid initialization (LCP 0.583 s), responsive interactivity (INP 30.5 ms), minimal layout shifts (CLS 0.01), and acceptable overhead (average 181.5 ms initial WASM module load). Although browser execution introduced performance penalties ( $$\sim $$ ∼ 130 $$\times $$ × slower than native), BioChef workflows still significantly outperformed traditional web services such as Galaxy by avoiding network delays and server-side queueing (11.3 $$\times $$ × faster in a standard pipeline benchmark). Conclusions BioChef shows how WebAssembly on the client side can democratize genomic data processing, ensuring privacy, reproducibility and ease of use without external dependencies. To our knowledge, this is the first fully client-side, graphical genomic workflow environment powered by WASM. Joaquim Vertentes Rosa, João Andrade, Jorge Miguel 0002, José Luís Oliveira |
BMC Bioinform. | 4 |
| 2025 | A Comparative Analysis of Ai-Based Solutions for Clinical DocumentationabstractHealthcare systems handle thousands of documents daily across various departments, requiring some effort during the digitalization processes. One strategy employed by medical staff is recording appointments for later transcription. However, this process is time-consuming and not practical for all scenarios. In this paper, we present a comprehensive methodology for converting medical audio recordings into structured documentation through multiple AI-based solutions. We propose and evaluate three distinct methods: a baseline two-stage pipeline using Mixtral 7B and Llama 70B models, a cyclic LLM approach leveraging self-improvement loops, and an embedding-based retrieval system utilizing BGE M3. Our experimental results show that the RBF kernel consistently outperformed linear kernels and logistic regression approaches across all metrics, maintaining high precision (0.87-0.94) and perfect recall. Luís Carlos Afonso, João Rafael Almeida, José Luís Oliveira |
CBMS | 3 |
| 2025 | An Embedding-Based Method for Processing Medical Audio into Structured ReportsabstractHealthcare professionals spend a significant portion of their time on electronic health records documentation, reducing patient interaction time and increasing operational costs. Our solution implements a two-stage pipeline combining voice-to-text transcription using WhisperV3 and text-to-structure conversion through an embedding-based approach. We address critical challenges in medical documentation automation, including specialized vocabulary processing and the prevention of hallucinations in generated content. The system was developed with continuous input from medical domain experts, resulting in a comprehensive field structure covering 78 essential information categories organized into six distinct sections. Our application processes clinical conversations locally, prioritizing data privacy and security while transforming unstructured medical notes into structured clinical documentation. The resulting system enables healthcare professionals to focus more time on patient care while simultaneously improving the quality and accessibility of medical records for better clinical decision-making. Luís Carlos Afonso, Carolina Gonçalves, Catarina Sousa, José Luís Oliveira |
CBMS | 4 |
| 2025 | An Embedding-Based Machine Learning Solution for Medical Concept MappingabstractThe integration of heterogeneous clinical datasets represents a fundamental challenge in contemporary biomedical research, particularly when reconciling multi-language and multi-institution data sources. The challenge of this procedure lies in the effort required to map the original concepts with their standard definitions. Various automated mapping solutions can assist researchers in this process, but the complexity grows when handling multi-language datasets, resulting in substantial manual work for translation and mapping. In this paper, we proposed a novel framework for clinical concept harmonisation that leverages vector-based embeddings and semantic search methodologies to enhance interoperability in multi-cohort studies. The methodology incorporates comprehensive data profiling, ontology-driven concept alignment, and machine learning-based vector search within a unified architecture. We demonstrate the efficacy of this approach through practical application to Alzheimer's disease (AD) research datasets from distinct institutions with different languages, achieving effective cross-lingual concept mapping while maintaining compatibility with established standardisation frameworks. Vicente Barros, Raquel Paradinha, João Rafael Almeida, José Luís Oliveira |
CBMS | 4 |
| 2025 | A Federated Random Forest Solution for Secure Distributed Machine LearningabstractPrivacy and regulatory barriers often hinder centralized machine learning solutions, particularly in sectors like healthcare where data cannot be freely shared. Federated learning has emerged as a powerful paradigm to address these concerns; however, existing frameworks primarily support gradientbased models, leaving a gap for more interpretable, tree-based approaches. This paper introduces a federated learning framework for Random Forest classifiers that preserves data privacy and provides robust performance in distributed settings. By leveraging PySyft for secure, privacy-aware computation, our method enables multiple institutions to collaboratively train Random Forest models on locally stored data without exposing sensitive information. The framework supports weighted model averaging to account for varying data distributions, incremental learning to progressively refine models, and local evaluation to assess performance across heterogeneous datasets. Experiments on two realworld healthcare benchmarks demonstrate that the federated approach maintains competitive predictive accuracy—within a maximum 9% margin of centralized methods—while satisfying stringent privacy requirements. These findings underscore the viability of tree-based federated learning for scenarios where data cannot be centralized due to regulatory, competitive, or technical constraints. The proposed solution addresses a notable gap in existing federated learning libraries, offering an adaptable tool for secure distributed machine learning tasks that demand both transparency and reliable performance. The tool is available at https://github.com/ieeta-pt/fed_rf Alexandre Cotorobai, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 3 |
| 2025 | Graph-based Optimization for Assembly Line Balancing Incorporating Metabolic RestrictionsabstractManual assembly processes remain a fundamental aspect of the manufacturing industry, primarily due to the dexterity and adaptability of human workers. However, the repetitive and physically demanding nature of these tasks highlights the need for an ergonomic and well-balanced workload, as poor ergonomics can also contribute to errors. This paper aims to automate the assembly line balancing process for a real-world case study. A new multi-objective Assembly Line Balancing Problem (ALBP) formulation is proposed, minimizing both workload variance and variance in workers’ energy expenditure across workstations while accounting for daily fluctuations in the number of operators. The proposed metabolic and time-sensitive assembly line balancing method integrates caloric considerations to reduce long-term risks of Work-Related Musculoskeletal Disorders (WMSDs). Using a graph-based approach, the framework ensures that precedence constraints are met while optimizing task assignments to minimize cycle time and metabolic cost per operator. To improve accuracy in estimating energy expenditure, this study employs the Methods-Time Measurement - Universal Analyzing System (MTM-UAS), which decomposes tasks into standardized motion elements. The methodology is validated through a real-world case study at Bosch Thermotechnology in Portugal, assigning tasks while minimizing the trade-offs between worker fatigue and production goals. The code to validate this study is publicly available at https://github.com/joaorafaelalmeida/line-balancing-algorithm. Joana Rafaela Almeida, Ana Moura, João Rafael Almeida, José Luís Oliveira |
CoDIT | 4 |
| 2025 | Securing DevOps by Identifying the Most Common Vulnerabilities in CI/CD PipelinesabstractSoftware engineering strategies have been studied over the last years, aiming to optimize the development of applications. Continuous Integration/Continuous Delivery (CI/CD) pipelines were a result of agile methodologies created to optimize those processes. While these pipelines can bring significant benefits to organizations, such as faster development and delivery, they also come with security risks. Such issues are sometimes ignored since the tools for automatically analyzing the security breaches in the application are focused on its artifacts (including source code, configurations, environments, and sandboxes), ignoring the CI/CD infrastructure. In this article, we explore the potential security flaws that may exist in CI/CD pipelines, along with the challenges and opportunities to investigate new solutions to mitigate such flaws. The article also proposes different models and strategies for protecting CI/CD pipelines. Luís Miguel Batista, Dinis Barroqueiro Cruz, Dimitri Silva, João Rafael Almeida, José Luís Oliveira |
ISCC | 5 |
| 2024 | Assessing the feasibility of observational data sources for multicenter clinical studiesabstractThe availability of large Electronic Health Records (EHR) databases has created new opportunities for clinical research. To improve interoperability of these databases a Common Data Model (CDM) can be used which enables standardized analytics. However, identifying the appropriate data sources for a specific study remains a challenge. The current strategy used for database discovery are based on catalogues that contain metadata which is often not rich enough for the task at hand. Additionally, sometimes this information is incorrect since it is inserted in such platforms manually by the data owners. In response to this challenge, we proposed the Concept Browser tool. The tool aims to streamline the process of efficiently exploring and selecting suitable OMOP CDM databases aligned with the study purpose. It was developed and validated in the EHDEN project, which currently contains information from more than 180 health databases across Europe and is now being used also in the DARWIN EU®initiative of the European Medicines Agency. João Rafael Almeida, Peter R. Rijnbeek, Maxim Moinat, José Luís Oliveira |
CBMS | 4 |
| 2024 | HealthDBFinder: a question-answering task for health database discoveryabstractIntegrating advanced data processing technologies into healthcare has shifted the medical studies paradigm. These evolve from data collection into management and analysis of Electronic Health Records (EHR) data. This change improved patient care and expanded the scope of clinical research through the secondary usage of existing data. Even though this problem was already solved in other initiatives, it raised new challenges, namely regarding cohort definition, data discovery, and evaluating the study feasibility. There are database catalogues to help in those tasks, but these fail in some cases due to insufficient information. Therefore, in this paper, we address this challenge by proposing a baseline method for information retrieval, including a synthetic dataset for further research. The information present in the dataset was generated from metadata extracted from real-world databases, which represents real problems that do not yet have a solution. The source code of this work is available at http://github.com/bioinformatics-ua/HealthDBFinder. João Rafael Almeida, Jorge Miguel 0002, Luís Carlos Afonso, Tiago Melo Almeida, Rui Antunes 0002, Richard Adolph Aires Jonker, João António Reis, Dimitri Alexandre da Silva, Sérgio Matos, José Luís Oliveira |
CBMS | 10 |
| 2024 | TAG-DTA: Binding-region-guided strategy to predict drug-target affinity using transformersabstractThe proper assessment of target-specific compound selectivity is paramount in the drug discovery context, promoting the identification of drug-target interactions (DTIs) and the discovery of potential leads. On that account, the accurate prediction of an unbiased drug-target binding affinity (DTA) metric is pivotal to understanding the binding process. Most in silico computational approaches, however, neglect the inter-dependency of the proteomics, chemical, and pharmacological spaces and the explainability during the model construction. Furthermore, these methods have yet to actively include information associated with binding pockets during the learning process, which is essential to DTA prediction performance and model explainability. In this study, we propose an end-to-end binding-region-guided Transformer-based architecture that simultaneously predicts the 1D binding pocket and the binding affinity of DTI pairs, where the prediction of the 1D binding pocket guides and conditions the prediction of DTA. This architecture uses 1D raw sequential and structural data to represent the proteins and compounds, respectively, and combines multiple Transformer-Encoder blocks to capture and learn the proteomics, chemical, and pharmacological contexts. The predicted 1D binding pocket conditions the attention mechanism of the Transformer-Encoder used to learn the pharmacological space in order to model the inter-dependency amongst binding-related positions. The results show that the proposed architecture, TAG-DTA, achieved the best performance in DTA prediction compared to state-of-the-art benchmarks, including in unknown subsets of the proteomics and chemical representation spaces. Moreover, the 1D binding pocket prediction increases the discriminative power and robustness of the aggregate representation of the pharmacological space and improves the DTA prediction performance. Overall, this research study validates the applicability of an end-to-end Transformer-based architecture in the context of drug discovery, and that combining computationally different yet contextually related tasks is critical to new findings in the DTI domain. Additionally, it shows that TAG-DTA is capable of providing increasing DTI and prediction understanding due to the nature of the attention blocks and prediction of the 1D binding pocket. The data and source code used in this study are available at: https://github.com/larngroup/TAG-DTA. Nelson R. C. Monteiro, José Luís Oliveira, Joel Arrais |
Expert Syst. Appl. | 2 |
| 2023 | A FAIR Approach to Real-World Health Data Management and AnalysisabstractThe increasing of health data sources to support clinical practice is opening the path for its secondary use in biomedical research. This changes the research paradigm, from data generation to data management and analysis. Although the potential for secondary use of this data is vast, including the improvement of healthcare systems and the advancement of clinical research, data discovery is challenging. In order to maximize data reusability, the FAIR principles have been developed as a guiding framework for system development. Nevertheless, the discovery and reuse of biomedical data present two main challenges: i) data partners grappling with ethical and social concerns related to data discoverability; ii) clinical researchers struggling to find the best data sources for their research studies. In this paper, we present a platform that provides a set of tools, compliant with the FAIR principles, to help data custodians when sharing data about biomedical databases, while allowing researchers to search for and select databases that meet their specific research needs. João Rafael Almeida, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 3 |
| 2023 | SecureFASTA: Ensuring privacy and trust when sharing genomic dataabstractGenomics has profoundly influenced the field of medicine, with advancements in DNA sequencing contributing to personalized medicine and a more comprehensive understanding of various diseases' genomic underpinnings. Sharing genomic data is vital for progressing the field and devising novel approaches to decipher the genome. Nevertheless, the sensitive nature of this information necessitates robust security measures for protection during storage and transfer. In this paper, we introduce SecureFASTA, a novel tool for securely encrypting and decrypting FASTA files without requiring a shared secret while minimizing the number of keys exchanged between pairs. Our approach combines symmetric and asymmetric encryption techniques, utilizing the Advanced Encryption Standard (AES) cypher and Rivest-Shamir-Adleman (RSA) encryption. Additionally, we implement a checksum function using the Secure Hash Algorithm (SHA-256) to verify the integrity of transferred FASTA files. Our evaluation demonstrates that SecureFASTA is fast, reliable, and secure, surpassing existing tools in terms of security and user-friendliness. Consequently, it offers a valuable solution for securely sharing and leveraging sensitive genomic data, marking a significant breakthrough in genomics. The tool's source code is available at https://github.com/bioinformatics-ua/SecureFASTA. Diniz Cruz, João Rafael Almeida, Jorge Miguel 0002, José Luís Oliveira |
CBMS | 4 |
| 2023 | A Multimodal Image Registration System for Histology ImagesabstractHistology image registration involves aligning microscopy images for various purposes, including creating 3D reconstructions from 2D, combining data from slices with different stain samples, and from multimodal registration. However, this process poses several challenges, including high resolution, non-linear elastic deformation, occlusions, missing sections, non-rigid deformation, contrast differences, and differences in appearance and local structure. Multimodal image registration is particularly challenging because different modalities may have different characteristics and require specific optimization algorithms. To address this issue, it is important to develop software that allows users to test different image registration algorithms and combine annotations made on them, in order to leverage the benefits of multiple modalities. To address this challenge, we developed a cloud-based Multimodal Image Registration system that enables developers and researchers to visually test the outcomes of various image registration algorithms. The system includes a project manager, an algorithm manager, and an image visualization system. The system was developed using the framework Django, JavaScript, and multiple libraries that facilitate the management and annotation of very high-resolution images. To demonstrate the effectiveness and flexibility of our system, we tested it using two different algorithms, SIFT and ORB, on nonlinear multimodal and brightfield images using the Hematoxylin and Eosin staining methods. The results show the system's ability to handle challenging image registration tasks while providing visualization tools to improve user experience. Rodrigo Escobar Díaz Guerrero, Yubraj Gupta, Thomas Bocklitz, José Luís Oliveira |
CBMS | 4 |
| 2023 | Querying semantic catalogues of biomedical databasesabstractBACKGROUND: Secondary use of health data is a valuable source of knowledge that boosts observational studies, leading to important discoveries in the medical and biomedical sciences. The fundamental guiding principle for performing a successful observational study is the research question and the approach in advance of executing a study. However, in multi-centre studies, finding suitable datasets to support the study is challenging, time-consuming, and sometimes impossible without a deep understanding of each dataset. METHODS: We propose a strategy for retrieving biomedical datasets of interest that were semantically annotated, using an interface built by applying a methodology for transforming natural language questions into formal language queries. The advantages of creating biomedical semantic data are enhanced by using natural language interfaces to issue complex queries without manipulating a logical query language. RESULTS: Our methodology was validated using Alzheimer's disease datasets published in a European platform for sharing and reusing biomedical data. We converted data to semantic information format using biomedical ontologies in everyday use in the biomedical community and published it as a FAIR endpoint. We have considered natural language questions of three types: single-concept questions, questions with exclusion criteria, and multi-concept questions. Finally, we analysed the performance of the question-answering module we used and its limitations. The source code is publicly available at https://bioinformatics-ua.github.io/BioKBQA/. CONCLUSION: We propose a strategy for using information extracted from biomedical data and transformed into a semantic format using open biomedical ontologies. Our method uses natural language to formulate questions to be answered by this semantic data without the direct use of formal query languages. Arnaldo Pereira, João Rafael Almeida, Rui Pedro Lopes, José Luís Oliveira |
J. Biomed. Informatics | 4 |
| 2022 | Portuguese Twitter Dataset on COVID-19abstractOver the last two years, the COVID-19 pandemic has affected hundreds of millions of people around the world. As in many crises, people turn to social media platforms, like Twitter, to communicate and share information. Twitter datasets have been used over the years in many research studies to extract valuable information. Therefore, several large COVID-19 Twitter datasets have been released over the last two years. However, none of these datasets contains only Portuguese Tweets, despite the Portuguese Language being reported as one of the top five languages used on Twitter. In this paper, we present the first large-scale Portuguese COVID-19 Twitter dataset. The dataset contains over 19 million Tweets spanning 2020 and 2021, allowing the entire pandemic to be analyzed. We also conducted a sentiment analysis on the dataset and correlated the various spikes in Tweet count and sentiment scores to various news articles and government announcements in Portugal and Brazil. The dataset is available at: https://github.com/bioinformatics-ua/Portuguese-Covid19-Dataset Richard Adolph Aires Jonker, Roshan Poudel, Olga Fajarda, Sérgio Matos, José Luís Oliveira, Rui Pedro Lopes |
ASONAM | 5 |
| 2022 | A secure architecture for exploring patient-level databases from distributed institutionsabstractOne of the main goals of clinical studies consists of identifying diseases' causes and improving the efficacy of medical treatments. Sometimes, the reduced number of participants is a limiting factor for these studies, leading researchers to organise multi-centre studies. However, sharing health data raises certain concerns regarding patients' privacy, namely related to the robustness of anonymisation procedures. Although these techniques remove personal identifiers from registries, some studies have shown that anonymisation procedures can sometimes be reverted using specific patients' characteristics. In this paper, we propose a secure architecture to explore distributed databases without compromising the patient's privacy. The proposed architecture is based on interoperable repositories supported by a common data model. João Rafael Almeida, João Paulo Barraca, José Luís Oliveira |
CBMS | 3 |
| 2022 | Combining heterogeneous patient-level data into tranSMART to support multicentre studiesabstractMany medical studies have been conducted aiming for better understanding of the causes of diseases and to assist in treatments and protective factors. In some cases, these studies do not produce impactful findings due to the small number of participants. Some initiatives already invested efforts in conducting multicentre studies, which raises other technical challenges due to the heterogeneity of datasets. The analysis of such data sources implies dealing with different data structures, terminologies, concepts, languages, and most importantly, the knowledge behind the data. In this paper, we present a methodology to centralise different datasets into the tranSMART application, using a harmonising strategy based on standard data schema. This methodology can help researchers to generate evidence from a wider variety of data sources. This proposal was validated using Alzheimer's Disease cohorts from several countries, combining at the end 6,669 subjects and 172 clinical concepts. The harmonised datasets can provide multi-cohort queries and analysis. The software package is available, under the MIT license, at https://github.com/bioinformatics-ua/tranSMART-migrator. João Rafael Almeida, Luís Bastião, Alejandro Pazos, José Luís Oliveira |
CBMS | 4 |
| 2022 | Visualising Time-evolving Semantic Biomedical DataabstractToday, medical studies enable a deeper understanding of health conditions, diseases and treatments, helping to improve medical care services. In observational studies, an adequate selection of datasets is important, to ensure the study's success and the quality of the results obtained. During the feasibility study phase, inclusion and exclusion criteria are defined, together with specific database characteristics to construct the cohort. However, it is not easy to compare database characteristics and their evolution over time during this selection. Data comparisons can be made using the data properties and aggregations, but the inclusion of temporal information becomes more complex due to the continuous evolution of concepts over time. In this paper, we propose two visualisation methods aiming for a better description of data evolution in clinical registers using biomedical standard vocabularies. Arnaldo Pereira, João Rafael Almeida, Rui Pedro Lopes, José Luís Oliveira |
CBMS | 4 |
| 2022 | Explainable deep drug-target representations for binding affinity predictionabstractBACKGROUND: Several computational advances have been achieved in the drug discovery field, promoting the identification of novel drug-target interactions and new leads. However, most of these methodologies have been overlooking the importance of providing explanations to the decision-making process of deep learning architectures. In this research study, we explore the reliability of convolutional neural networks (CNNs) at identifying relevant regions for binding, specifically binding sites and motifs, and the significance of the deep representations extracted by providing explanations to the model's decisions based on the identification of the input regions that contributed the most to the prediction. We make use of an end-to-end deep learning architecture to predict binding affinity, where CNNs are exploited in their capacity to automatically identify and extract discriminating deep representations from 1D sequential and structural data. RESULTS: The results demonstrate the effectiveness of the deep representations extracted from CNNs in the prediction of drug-target interactions. CNNs were found to identify and extract features from regions relevant for the interaction, where the weight associated with these spots was in the range of those with the highest positive influence given by the CNNs in the prediction. The end-to-end deep learning model achieved the highest performance both in the prediction of the binding affinity and on the ability to correctly distinguish the interaction strength rank order when compared to baseline approaches. CONCLUSIONS: This research study validates the potential applicability of an end-to-end deep learning architecture in the context of drug discovery beyond the confined space of proteins and ligands with determined 3D structure. Furthermore, it shows the reliability of the deep representations extracted from the CNNs by providing explainability to the decision-making process. Nelson R. C. Monteiro, Carlos J. V. Simões, Henrique V. Ávila, Maryam Abbasi, José Luís Oliveira, Joel Arrais |
BMC Bioinform. | 5 |
| 2022 | Systematic review of question answering over knowledge basesabstractAbstract Over the years, a growing number of semantic data repositories have been made available on the web. However, this has created new challenges in exploiting these resources efficiently. Querying services require knowledge beyond the typical user’s expertise, which is a critical issue in adopting semantic information solutions. Several proposals to overcome this difficulty have suggested using question answering (QA) systems to provide user‐friendly interfaces and allow natural language use. Because question answering over knowledge bases (KBQAs) is a very active research topic, a comprehensive view of the field is essential. The purpose of this study was to conduct a systematic review of methods and systems for KBQAs to identify their main advantages and limitations. The inclusion criteria rationale was English full‐text articles published since 2015 on methods and systems for KBQAs. Sixty‐six articles were reviewed to describe their underlying reference architectures. Arnaldo Pereira, Alina Trifan, Rui Pedro Lopes, José Luís Oliveira |
IET Softw. | 4 |
| 2021 | An Architecture to Define Cohorts over Medical Imaging DatasetsabstractThe DICOM standard has been widely adopted for the exchange and management of biomedical images. Its hierarchical structure allows representing data and metadata of medical imaging studies. However, other patient data not directly related with the study, such as prescriptions and treatments, are stored in independent Electronic Health Record (EHR) systems. With the increasing production of medical imaging studies, repositories responsible for storing DICOM images started to contain massive amounts of data. Therefore, retrieving a subset of images based on similar criteria as the ones used in EHR systems, is a complex task for a medical researcher. In this paper, we propose an architecture to define cohorts over medical imaging data sets. This proposal uses a DICOM archive to index and to retrieve images, while the studies' selection is performed through a web application, ATLAS, which is normally used on observational studies upon EHR data. The presented architecture was validated using a public data set with synthetic EHR data. João Rafael Almeida, Eriksson J. Melicio Monteiro, José Luís Oliveira |
CBMS | 3 |
| 2021 | Improvements in lymphocytes detection using deep learning with a preprocessing stageabstractLymphocytes are a type of white blood cell that are part of the adaptive immune system and respond to infectious microorganisms. Due to this key role, its detection and quantification allow analyzing the overall status of the immune system. However, the manual detection of lymphocytes in tissue slices is a laborious task, and it depends on the expertise of the observer, reason why an automated image analysis helps to speedup this process. Several different techniques have been used to automatize this task, such as morphological operations, classification algorithms, and, more recently, deep learning approaches. In this work, we propose two preprocessing methods for improving the lymphocytes detection in digital images. Furthermore, this study proposes a change in the ground truth, in order to turn it into a segmentation map, and evaluate semantic segmentation models in a dataset that originally does not allow this approach. Two deep learning models (Segnet and U-Net) with different backbones (VGG16 and Resnet50) were used for the training and test sets. One of the proposed methods showed an F1-score 11% higher than simply using a color normalization. The results were compared with other state-of-the-art studies, showing one of the best-ranked results. Rodrigo Escobar Díaz Guerrero, José Luís Oliveira |
CBMS | 2 |
| 2021 | Easing the Questioning of Semantic Biomedical DataabstractResearchers have been using semantic technologies as essential tools to structure knowledge. This is particularly relevant in the biomedical domain, where large dataset are continuously generated. Semantic technologies offer the ability to describe data and to map and linking distributed repositories, creating a network where the searching interface is a single entry point. However, the increasing number of semantic data repositories that are publicly available is creating new challenges related to its exploration. Despite being human and machine-readable, these technologies are much more challenging for end-users. Querying services usually require mastering formal languages and that knowledge is beyond the typical user's expertise, being a critical issue in adopting semantic web information systems. In particular, the questioning of biomedical data presents specific challenges for which there are still no mature proposals for production environments. This paper presents a solution to query biomedical semantic databases using natural language. The system is at the intersection between semantic parsing and the use of templates. It makes it possible to extract information in a friendly way for users who are not experts in semantic queries. Arnaldo Pereira, Rui Pedro Lopes, José Luís Oliveira |
CBMS | 3 |
| 2021 | A Comparative Analysis of Data Platforms for Rare DiseasesabstractThe increasing interest in finding drugs and treatments for rare diseases led to the creation of research studies and clinical trials which data and results have been stored in multiple, heterogeneous databases. The lack of data harmonisation, combined with the need to improve current medical knowledge, has encouraged the research community to create computational solutions to aggregate this information. Although such platforms were created in the same area, orphan diseases, they were normally developed for different purposes, increasing the task complexity for end-users when needing to search for gene-to-phenotype information (e.g. genes, mutations, symptoms, etc.). Aiming to help answer these questions, we conducted a comprehensive analysis of the existent platforms designed to retrieve and visualise information about genetic rare diseases. In this analysis, we found several platforms from which we identified 7 candidates based on a set of inclusion and exclusion criteria. Through this analysis we were able to assess each system's characteristics and identify the most appropriate for distinct use cases and audiences, namely medical researchers, bioinformaticians and patients and relatives. Mariana Sequeira, João Rafael Almeida, José Luís Oliveira |
CBMS | 3 |
| 2021 | Optimizing blood-brain barrier permeation through deep reinforcement learning for de novo drug designabstractMOTIVATION: The process of placing new drugs into the market is time-consuming, expensive and complex. The application of computational methods for designing molecules with bespoke properties can contribute to saving resources throughout this process. However, the fundamental properties to be optimized are often not considered or conflicting with each other. In this work, we propose a novel approach to consider both the biological property and the bioavailability of compounds through a deep reinforcement learning framework for the targeted generation of compounds. We aim to obtain a promising set of selective compounds for the adenosine A2A receptor and, simultaneously, that have the necessary properties in terms of solubility and permeability across the blood-brain barrier to reach the site of action. The cornerstone of the framework is based on a recurrent neural network architecture, the Generator. It seeks to learn the building rules of valid molecules to sample new compounds further. Also, two Predictors are trained to estimate the properties of interest of the new molecules. Finally, the fine-tuning of the Generator was performed with reinforcement learning, integrated with multi-objective optimization and exploratory techniques to ensure that the Generator is adequately biased. RESULTS: The biased Generator can generate an interesting set of molecules, with approximately 85% having the two fundamental properties biased as desired. Thus, this approach has transformed a general molecule generator into a model focused on optimizing specific objectives. Furthermore, the molecules' synthesizability and drug-likeness demonstrate the potential applicability of the de novo drug design in medicinal chemistry. AVAILABILITY AND IMPLEMENTATION: All code is publicly available in the https://github.com/larngroup/De-Novo-Drug-Design. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tiago Pereira 0001, Maryam Abbasi, José Luís Oliveira, Bernardete Ribeiro, Joel Arrais |
Bioinform. | 3 |
| 2021 | A two-stage workflow to extract and harmonize drug mentions from clinical notes into observational databases
João Rafael Almeida, João Figueira Silva, Sérgio Matos, José Luís Oliveira |
J. Biomed. Informatics | 4 |
| 2020 | Exploring a Siamese Neural Network Architecture for One-Shot Drug DiscoveryabstractThe application of deep neural networks in drug discovery is mainly due to their enormous potential to significantly increase the predictive power when inferring the properties and activities of small-molecules. However, in the traditional drug discovery process, where supervised data is scarce, the lead-optimization step is a low-data problem, making it difficult to find molecules with the desired therapeutic activity and obtain accurate predictions for candidate compounds. One major requirement to ensure the validity of the obtained neural network models is the need for a large number of training examples per class, which is not always feasible in drug discovery applications. This invalidates the use of instances whose classes were not considered in the training phase or in data where the number of classes is high and oscillates dynamically. The main objective of the study is to optimize the discovery of novel compounds based on a reduced set of candidate drugs. We propose a Siamese neural network architecture for one-shot classification, based on Convolutional Neural Networks (CNNs), that learns from a similarity score between two input molecules according to a given similarity function. Using a one-shot learning strategy, few instances per class are needed for training, and a small amount of data and computational resources are required to build an accurate model. The results achieved demonstrate that using a Siamese Deep Neural Network for one-shot classification leads to overall improved performance when compared to other state-of the-art models. The proposed architecture provides an accurate and reliable prediction of novel compounds considering the lack of biological data available for drug discovery tasks. Luis H. M. Torres, Nelson R. C. Monteiro, José Luís Oliveira, Joel Arrais, Bernardete Ribeiro |
BIBE | 3 |
| 2020 | A Recommender System to Help Discovering Cohorts in Rare DiseasesabstractCohort studies have been playing a key role in helping our understanding of diseases, health conditions, and treatments. These cohorts are often composed of a small number of subjects, especially in rare diseases studies, which reduces the statistical power of the results. One solution that can strengthen the scientific findings is to combine distinct studies and perform then multi-cohort analysis. However, even studies conducted for the same purpose in distinct research groups can have different scopes and medical observations, which preclude across-cohort exploration. In this paper, we propose a recommendation system to automatically discover cohorts of interest. This methodology uses context-based retrieval techniques combined with collaborative filtering to find relevant cohorts and scientific literature about a specific clinical investigation. The system was validated in a community focused on the study of Alzheimer's diseases, which includes 62 cohorts. João Rafael Almeida, Eriksson J. Melicio Monteiro, Luís Bastião, Alejandro Pazos, José Luís Oliveira |
CBMS | 5 |
| 2020 | Multi-language Concept Normalisation of Clinical CohortsabstractThe exploration of multiple cohorts allows researchers to answer new research questions using more substantial clinical data. However, this is only possible if the cohorts are interoperable, which implies the migration of the original cohort into a common data schema. The problem of this procedure is the effort necessary to map the original concepts into their standard definitions. While several automatic mapping solutions can help in this task, its complexity increases when dealing with multi-language cohorts, leading to a significant manual effort in translating and mapping. In this paper we propose a system that combines text mining with language detection techniques, aiming to optimise these migration pipelines. This system was designed to be integrated into already existing migration workflows, without the need of adapting them. The system was validated using Alzheimer's diseases cohorts, but it is enough general to be applied in other use cases. João Rafael Almeida, José Luís Oliveira |
CBMS | 2 |
| 2020 | Understanding Depression from Psycholinguistic Patterns in Social Media Texts
Alina Trifan, Rui Antunes 0002, Sérgio Matos, José Luís Oliveira |
ECIR (2) | 4 |
| 2019 | GenericCDSS - A Generic Clinical Decision Support SystemabstractClinical decision support systems (CDSS) are currently essential tools to guide medical diagnostics and patients' treatments, and they are specially important for the better care management of chronic diseases, such as cancer and diabetes. These systems help to decide on the best treatment solution, namely in centres where there is a shortage of medical experts. CDSS tools are often integrated into the Electronic Health Record (EHR) to facilitate the reuse of patient data. However, many times, creating new and intuitive protocols that are disease-specific is still a challenge. In this paper we present an open source solution (GenericCDSS) that can be used to streamline the development of autonomous CDSS, avoiding the dependency on third-party tools to manage patient data and clinical protocols. The software tool provides a modern user interface, supporting multi-platforms such as mobile and desktop devices. GenericCDSS is publicly available at https://github.com/bioinformatics-ua/GenericCDSS, under a GNU GPL license. João Rafael Almeida, José Luís Oliveira |
CBMS | 2 |
| 2019 | Image selection based on low level properties for lifelog moment retrievalabstractThe increasing number of mobile and wearable devices is dramatically changing the way we collect data about person’s life. These devices allow recording our daily activities and behavior in several forms, e.g., text, images, bio-signals, or video. However, many times, the collected data includes low quality or irrelevant contents, feeding lifelogging applications with huge amounts of data, and creating computational challenges for patterns’ identification. In this paper, we propose a fast image analysis approach to automatically select relevant images from lifelog data. Using images intrinsic information, such as scenes and objects, we have manually curated two datasets, one with relevant content and another with non-relevant information. Then, we applied supervised learning algorithms based on low-level image features, namely blur and focus, to find the binary model that best discriminates between the two classes. The binary models were then compared based on learning curves and f1-scores, achieving a 95.4% of f1-score for the best one. By reducing the amount of images in the lifelog data, we were able to save computational time without losing images with relevant content. Ricardo F. Ribeiro, António J. R. Neves, José Luís Oliveira |
ICMV | 3 |
| 2019 | Patient data discovery platforms as enablers of biomedical and translational research: A systematic reviewabstractBACKGROUND: The global shift from paper health records to electronic ones has led to an impressive growth of biomedical digital data along the past two decades. Exploring and extracting knowledge from these data has the potential to enhance translational research and lead to positive outcomes for the population's health and healthcare. OBECTIVE: The aim of this study was to conduct a systematic review to identify software platforms that enable discovery, secondary use and interoperability of biomedical data. Additionally, we aim evaluating the identified solutions in terms of clinical interest and main healthcare-related outcomes. METHODS: A systematic search of the scientific literature published and indexed in Pubmed between January 2014 and September 2018 was performed. Inclusion criteria were as follows: relevance for the topic of biomedical data discovery, English language, and free full text. To increase the recall, we developed a semi-automatic and incremental methodology to retrieve articles that cite one or more of the previous set. RESULTS: A total number of 500 candidate papers were retrieved through this methodology. Of these, 85 were eligible for abstract assessment. Finally, 37 studies qualified for a full-text review, and 20 provided enough information for the study objectives. CONCLUSIONS: This study revealed that biomedical discovery platforms are both a current necessity and a significantly innovative agent in the area of healthcare. The outcomes that were identified, in terms of scientific publications, clinical studies and research collaborations stand as evidence. Alina Trifan, José Luís Oliveira |
J. Biomed. Informatics | 2 |
| 2018 | Simplifying the Digitization of Clinical Protocols for Diabetes ManagementabstractHyperglycemia is a health condition characterized by abnormally high blood glucose, typically caused by a deficient usage, or lack, of insulin. Due to the metabolic derangements of this clinical condition, its regular monitoring, as well the administration of the most effective treatment, are major concerns for healthcare institutions. In this paper, we present a computational solution to build diabetes management protocols, which helps health professionals providing an adequate treatment for each hyperglycemic inpatient. João Rafael Almeida, Joana Guimaraes, José Luís Oliveira |
CBMS | 3 |
| 2018 | Services Orchestration and Workflow Management in Distributed Medical Imaging EnvironmentsabstractMedical imaging laboratories are supported by information and communication systems commonly denominated as PACS, that encompasses technology for acquisition, archive, distribution and visualization of digital images in network. Concerning the data and workflow management, traditional solutions used in production provide a limited set of services usually configured at system installation. As result, healthcare institutions are not able to fully explore their infrastructure or adapt it to new operational requirements, either for clinical or research procedures. This article proposes a framework for services orchestration and workflow management in distributed medical imaging environments. It was designed for end-user usage and is accessible through a Web portal that allows to document, repeat and allocate procedures and tasks to correct resources, either from information systems or human interventions. It provides an abstraction layer for integration with distinct data sources through standard services, allows the creation of new services through orchestration of existent ones and the scheduling of tasks. Moreover, it includes a logging and alert mechanism integrated with email service. The solution was validated through the specification of two use cases that were deployed in production environment. João Rafael Almeida, Tiago Marques Godinho, Luís Bastião, Carlos Costa 0001, José Luís Oliveira |
CBMS | 5 |
| 2018 | A FAIR Marketplace for Biomedical Data Custodians and Clinical ResearchersabstractThe exponential growth in the volume of biomedical data, induced by the increasing digitalization of health care services, has led to a shift of the bottleneck of biomedical research: from data generation to data management and analysis. When shared, the opportunities of secondary use for these data are end-less, from advancing clinical research to an overall improvement of healthcare systems. The FAIR principles have been designed as a guideline for the development of systems that enhance data reusability. The challenges of biomedical data discovery and reuse can be understood and addressed from two distinct points of view: on one side, data custodians are apprehensive about ethical and social issues when turning data discoverable; on the other one, clinical researchers have to perform intensive searches over geographical scattered, heterogeneous data sources and pursue extensive protocols for reusing these data. In this paper we present a web platform intended to serve as a bridge between data custodians and biomedical researchers. Data custodians are able to publish and share several levels of information about biomedical databases, while researchers can search for databases that fulfil their research requirements. Alina Trifan, José Luís Oliveira |
CBMS | 2 |
| 2018 | Simplifying biomedical data sharing through a web portal generatorabstractDespite the huge amount of biomedical data that is currently being collected, they end up many times neglected in internal silos, hindering its discovery and exploration at a large scale. One issue that contributes to this scenario is the complexity in constructing a full-fleshed web-system, to capture and manage these multiple and heterogeneous datasets. To help simplifying this task, we propose an semi-automated solution for the development of web-based data catalogs, targeted for non technical users. Starting from a simple data model description, the system automatically generates user interfaces and services for data management. This engine is being successfully used in several European research communities. Andre Malta, José Luís Oliveira |
HealthCom | 2 |
| 2016 | Ensemble-Based Methodology for the Prediction of Drug-Target InteractionsabstractAntibacterial resistance has been progressively increasing mostly due to selective antibiotic pressure, forcing pathogens to either adapt or die. The development of antibacterial resistance to last-line antibiotics urges the formulation of alternative strategies for drug discovery. Recently, attention has been devoted to the development of computational methods to predict drug-target interactions (DTIs). Here we present a computational strategy to predict proteome-scale DTIs based on the combination of the drugs' chemical features and substructural fingerprints, and on the structural information and physicochemical properties of the proteins. We propose an ensemble learning combination of Support-Vector Machine and Random Forest to deal with the complexity of DTI classification. Two distinct classification models were developed to ascertain whether taking the type of protein target (i.e., enzymes, g-protein-coupled receptors, ion channels and nuclear receptors) into account improves classification performance. External validation analysis was consistent with internal five-fold cross-validation, with an AUC of 0.87. This strategy was applied to the proteome of methicillin-resistant Staphylococcus aureus COL (MRSA COL, taxonomy id: 93062), a major nosocomial pathogen worldwide whose antimicrobial resistance and incidence rate keeps steadily increasing. Our predictive framework is available at http://bioinformatics.ua.pt/software/dtipred. Edgar D. Coelho, Joel Arrais, José Luís Oliveira |
CBMS | 3 |
| 2016 | Caching and Prefetching Images in a Web-Based DICOM ViewerabstractThe general trend of information access anywhere and anytime is also leading to the emergence of innovative medical imaging systems, adapted to this new reality. The HTML5 standard led Web applications to another software level, providing a set of features that allows developing professional Web-based medical imaging applications. However, despite the visualization quality that is already possible in HTML5 browsers, the performance is still an issue, due to the typical size of image studies. In this paper, we present a caching and prefetching solution that enriches the end-user experience in a Web-based DICOM viewer, by reducing data access latency of examinations under revision. We deployed the system in a radiology center for mammography screening, at a national level, and the results show that this technique significantly reduces the average examination access latency, during the reviewing process. Eriksson J. Melicio Monteiro, Carlos Costa 0001, José Luís Oliveira, David Campos 0001, Luís Bastião |
CBMS | 3 |
| 2016 | Corrigendum: EuGene: maximizing synthetic gene design for heterologous expressionabstractBioinformatics (2012) 28(20), 2683–2684 The authors of the above article wish to amend the funding acknowledgement, so that it reads as follows: This article was partially supported by the European projects MEPHITIS and GEN2PHEN and by FEDER funds through the Programa Operacional Factores de Competitividade - COMPETE (FCOMP-01-0124-FEDER-014284) and by National Funds through FCT - Foundation for Science and Technology in the scope of the research project ref. PTDC/BIA-GEN/110383/2009. P.G. was supported by FCT-SFRH/BD/71063/2010. The authors apologize for this error. Paulo Gaspar, José Luís Oliveira, Jörg Frommlet, Manuel A. S. Santos, Gabriela R. Moura |
Bioinform. | 2 |
| 2016 | Computational Discovery of Putative Leads for Drug Repositioning through Drug-Target Interaction PredictionabstractDe novo experimental drug discovery is an expensive and time-consuming task. It requires the identification of drug-target interactions (DTIs) towards targets of biological interest, either to inhibit or enhance a specific molecular function. Dedicated computational models for protein simulation and DTI prediction are crucial for speed and to reduce the costs associated with DTI identification. In this paper we present a computational pipeline that enables the discovery of putative leads for drug repositioning that can be applied to any microbial proteome, as long as the interactome of interest is at least partially known. Network metrics calculated for the interactome of the bacterial organism of interest were used to identify putative drug-targets. Then, a random forest classification model for DTI prediction was constructed using known DTI data from publicly available databases, resulting in an area under the ROC curve of 0.91 for classification of out-of-sampling data. A drug-target network was created by combining 3,081 unique ligands and the expected ten best drug targets. This network was used to predict new DTIs and to calculate the probability of the positive class, allowing the scoring of the predicted instances. Molecular docking experiments were performed on the best scoring DTI pairs and the results were compared with those of the same ligands with their original targets. The results obtained suggest that the proposed pipeline can be used in the identification of new leads for drug repositioning. The proposed classification model is available at http://bioinformatics.ua.pt/software/dtipred/. Edgar D. Coelho, Joel Arrais, José Luís Oliveira |
PLoS Comput. Biol. | 3 |
| 2015 | Ann2RDF: moving annotations to semantic webabstractThe annotation of concepts and susceptible interactions has been assuming a key role in the extraction of relevant information from published documents. However, distinct annotation tools generate also different formats, creating a barrier to efficiently combine and exchange this information. The migration of curated information into semantic web format and services provides an additional value to share that knowledge, but data transformation represents here an additional challenge. In this manuscript, we present a unified layer between text-mining tools and semantic web services to reduce the effort of combining different formats. The Ann2RDF is focused on reusing existing curated data from external text-mining tools to improve their availability through an open representation model. This result in a more suitable transition process, in which desired annotations are enriched with the possibility to be shared, compared and reused across semantic Knowledge Bases. Pedro Sernadela, Sérgio Matos, José Luís Oliveira |
iiWAS | 3 |
| 2015 | An automated real-time integration and interoperability framework for bioinformaticsabstractBACKGROUND: In recent years data integration has become an everyday undertaking for life sciences researchers. Aggregating and processing data from disparate sources, whether through specific developed software or via manual processes, is a common task for scientists. However, the scope and usability of the majority of current integration tools fail to deal with the fast growing and highly dynamic nature of biomedical data. RESULTS: In this work we introduce a reactive and event-driven framework that simplifies real-time data integration and interoperability. This platform facilitates otherwise difficult tasks, such as connecting heterogeneous services, indexing, linking and transferring data from distinct resources, or subscribing to notifications regarding the timeliness of dynamic data. For developers, the framework automates the deployment of integrative and interoperable bioinformatics applications, using atomic data storage for content change detection, and enabling agent-based intelligent extract, transform and load tasks. CONCLUSIONS: This work bridges the gap between the growing number of services, accessing specific data sources or algorithms, and the growing number of users, performing simple integration tasks on a recurring basis, through a streamlined workspace available to researchers and developers alike. Pedro Lopes 0002, José Luís Oliveira |
BMC Bioinform. | 2 |
| 2014 | geneCommittee: a web-based tool for extensively testing the discriminatory power of biologically relevant gene sets in microarray data classificationabstractBACKGROUND: The diagnosis and prognosis of several diseases can be shortened through the use of different large-scale genome experiments. In this context, microarrays can generate expression data for a huge set of genes. However, to obtain solid statistical evidence from the resulting data, it is necessary to train and to validate many classification techniques in order to find the best discriminative method. This is a time-consuming process that normally depends on intricate statistical tools. RESULTS: geneCommittee is a web-based interactive tool for routinely evaluating the discriminative classification power of custom hypothesis in the form of biologically relevant gene sets. While the user can work with different gene set collections and several microarray data files to configure specific classification experiments, the tool is able to run several tests in parallel. Provided with a straightforward and intuitive interface, geneCommittee is able to render valuable information for diagnostic analyses and clinical management decisions based on systematically evaluating custom hypothesis over different data sets using complementary classifiers, a key aspect in clinical research. CONCLUSIONS: geneCommittee allows the enrichment of microarrays raw data with gene functional annotations, producing integrated datasets that simplify the construction of better discriminative hypothesis, and allows the creation of a set of complementary classifiers. The trained committees can then be used for clinical research and diagnosis. Full documentation including common use cases and guided analysis workflows is freely available at http://sing.ei.uvigo.es/GC/. Miguel Reboiro-Jato, Joel Arrais, José Luís Oliveira, Florentino Fernández Riverola |
BMC Bioinform. | 3 |
| 2014 | XDS-I Outsourcing Proxy: Ensuring Confidentiality While Preserving InteroperabilityabstractThe interoperability of services and the sharing of health data have been a continuous goal for health professionals, patients, institutions, and policy makers. However, several issues have been hindering this goal, such as incompatible implementations of standards (e.g., HL7, DICOM), multiple ontologies, and security constraints. Cross-enterprise document sharing (XDS) workflows were proposed by Integrating the Healthcare Enterprise (IHE) to address current limitations in exchanging clinical data among organizations. To ensure data protection, XDS actors must be placed in trustworthy domains, which are normally inside such institutions. However, due to rapidly growing IT requirements, the outsourcing of resources in the Cloud is becoming very appealing. This paper presents a software proxy that enables the outsourcing of XDS architectural parts while preserving the interoperability, confidentiality, and searchability of clinical information. A key component in our architecture is a new searchable encryption (SE) scheme-Posterior Playfair Searchable Encryption (PPSE)-which, besides keeping the same confidentiality levels of the stored data, hides the search patterns to the adversary, bringing improvements when compared to the remaining practical state-of-the-art SE schemes. Luís S. Ribeiro, Carlos Viana-Ferreira, José Luís Oliveira, Carlos Costa 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2013 | Integrating echocardiogram reports with medical imagingabstractHealthcare institutions are increasingly taking advantage of information and computational systems to enhance the efficiency and quality of their services. These IT systems are normally able to handle huge amounts of digital data, and to extract relevant fingerprints useful to improve the quality of clinical practice. However, building automatized processes to achieve this over multiple and heterogeneous databases is still a challenge. This paper presents a new approach able to collect and index information from distinct medical data sources, allowing us to identify important metrics to evaluate the performance and quality of clinical services. A case study combining information from ultrasound medical images and echocardiogram clinical reports is also presented. Luís Bastião, Samuel Campos, Carlos Costa 0001, José Luís Oliveira |
CBMS | 4 |
| 2013 | A multi-domain platform for medical imagingabstractThe increasing adoption of medical imaging equipment in healthcare has been leading to a huge dispersion of data repositories and institutions. Although the quality of diagnostic and treatment is deeply dependent on the health information that is available for physicians, several legal and technological issues have hindered the integration of these data. One of such problems is because traditional medical imaging protocols do not perform well in inter-institutional scenarios. This paper describes a hybrid network platform for medical imaging systems that provides searching and retrieval over multiple centres. Three key components support the system: an indexing engine, a multicast framework and a cloud service. Using a peer-to-peer paradigm with security constraints, the platform gathers the information of medical imaging repositories hosted inside the institutions, allowing physicians to access data when and where they need it. Carlos Viana-Ferreira, Carlos Costa 0001, José Luís Oliveira |
CBMS | 3 |
| 2013 | A DICOM viewer based on web technologyabstractDuring the last decade, medical imaging services have assumed a central role in healthcare institutions, and they are nowadays a decisive factor for the quality of diagnostic and treatments. Health stakeholders and policy makers have been steadily adopting PACS and DICOM standard, simplifying interoperability between distinct equipment and institutions. To assist images' interpretation, several visualization solutions emerged. However, these applications are targeted to specific operating systems, hindering its ubiquitous use, in increasing web-based working environments. In this paper we present a Web-based DICOM viewer that was entirely developed with web technology, namely HTML5 and JavaScript. The result is a visualization station that is already in use in two medical imaging centres and that can be accessed through a common web browser, from any computer, mobile device, or operation system. Eriksson J. Melicio Monteiro, Carlos Costa 0001, José Luís Oliveira |
Healthcom | 3 |
| 2013 | Leveraging XDS-I and PIX workflows for validating cross-enterprise patient identity linkageabstractDocument exchange communities set the ground for cross-organization cooperation. They enable the exchange of patient's documents across distinct health organizations. However, there are various challenges that must be overcame before deploying such communities, for instance the construction of the Enterprise Master Patient Index (EMPI) which maps the several patient identifiers of each domain. This paper describes the development of an interoperable distributed system that expedites the exchange of documents by taking care of the patient identities autonomously. The system automatically builds the EMPI leveraging the healthcare workflow (based on PIX and XDS-I) for validating the automatic linkages of the patient identifiers. The human validation is a consequence of user's interaction with cross-domain documents: distributing and attenuating the validation effort. Luís S. Ribeiro, Frederico Honorio, José Luís Oliveira, Carlos Costa 0001 |
Healthcom | 3 |
| 2013 | Exploring nanopublications integration in pharmacovigilance scenariosabstractNowadays, the post-market assessment and monitoring of interactions amongst prescribed drugs and adverse events in the pharmacovigilance domain reveals previously unknown interactions. Involved stakeholders are realising that modern approaches, based on integrative data exploration environments, are vital to improve current drug research and development processes. In this manuscript we introduce a comprehensive pipeline for integrating large collections of annotated data, abstracting it to a Semantic Web environment, modelling it according to the nanopublications standard, and delivering straightforward access to collected knowledge. Pedro Sernadela, Pedro Lopes 0002, José Luís Oliveira |
Healthcom | 3 |
| 2013 | BeCAS: biomedical concept recognition services and visualizationabstractSUMMARY: The continuous growth of the biomedical scientific literature has been motivating the development of text-mining tools able to efficiently process all this information. Although numerous domain-specific solutions are available, there is no web-based concept-recognition system that combines the ability to select multiple concept types to annotate, to reference external databases and to automatically annotate nested and intercepted concepts. BeCAS, the Biomedical Concept Annotation System, is an API for biomedical concept identification and a web-based tool that addresses these limitations. MEDLINE abstracts or free text can be annotated directly in the web interface, where identified concepts are enriched with links to reference databases. Using its customizable widget, it can also be used to augment external web pages with concept highlighting features. Furthermore, all text-processing and annotation features are made available through an HTTP REST API, allowing integration in any text-processing pipeline. AVAILABILITY: BeCAS is freely available for non-commercial use at http://bioinformatics.ua.pt/becas. CONTACTS: [email protected] or [email protected]. Tiago Nunes, David Campos 0001, Sérgio Matos, José Luís Oliveira |
Bioinform. | 4 |
| 2013 | Gimli: open source and high-performance biomedical name recognitionabstractBACKGROUND: Automatic recognition of biomedical names is an essential task in biomedical information extraction, presenting several complex and unsolved challenges. In recent years, various solutions have been implemented to tackle this problem. However, limitations regarding system characteristics, customization and usability still hinder their wider application outside text mining research. RESULTS: We present Gimli, an open-source, state-of-the-art tool for automatic recognition of biomedical names. Gimli includes an extended set of implemented and user-selectable features, such as orthographic, morphological, linguistic-based, conjunctions and dictionary-based. A simple and fast method to combine different trained models is also provided. Gimli achieves an F-measure of 87.17% on GENETAG and 72.23% on JNLPBA corpus, significantly outperforming existing open-source solutions. CONCLUSIONS: Gimli is an off-the-shelf, ready to use tool for named-entity recognition, providing trained and optimized models for recognition of biomedical entities from scientific text. It can be used as a command line tool, offering full functionality, including training of new models and customization of the feature set and model parameters through a configuration file. Advanced users can integrate Gimli in their text mining workflows through the provided library, and extend or adapt its functionalities. Based on the underlying system characteristics and functionality, both for final users and developers, and on the reported performance results, we believe that Gimli is a state-of-the-art solution for biomedical NER, contributing to faster and better research in the field. Gimli is freely available at http://bioinformatics.ua.pt/gimli. David Campos 0001, Sérgio Matos, José Luís Oliveira |
BMC Bioinform. | 3 |
| 2013 | A modular framework for biomedical concept recognitionabstractBACKGROUND: Concept recognition is an essential task in biomedical information extraction, presenting several complex and unsolved challenges. The development of such solutions is typically performed in an ad-hoc manner or using general information extraction frameworks, which are not optimized for the biomedical domain and normally require the integration of complex external libraries and/or the development of custom tools. RESULTS: This article presents Neji, an open source framework optimized for biomedical concept recognition built around four key characteristics: modularity, scalability, speed, and usability. It integrates modules for biomedical natural language processing, such as sentence splitting, tokenization, lemmatization, part-of-speech tagging, chunking and dependency parsing. Concept recognition is provided through dictionary matching and machine learning with normalization methods. Neji also integrates an innovative concept tree implementation, supporting overlapped concept names and respective disambiguation techniques. The most popular input and output formats, namely Pubmed XML, IeXML, CoNLL and A1, are also supported. On top of the built-in functionalities, developers and researchers can implement new processing modules or pipelines, or use the provided command-line interface tool to build their own solutions, applying the most appropriate techniques to identify heterogeneous biomedical concepts. Neji was evaluated against three gold standard corpora with heterogeneous biomedical concepts (CRAFT, AnEM and NCBI disease corpus), achieving high performance results on named entity recognition (F1-measure for overlap matching: species 95%, cell 92%, cellular components 83%, gene and proteins 76%, chemicals 65%, biological processes and molecular functions 63%, disorders 85%, and anatomical entities 82%) and on entity normalization (F1-measure for overlap name matching and correct identifier included in the returned list of identifiers: species 88%, cell 71%, cellular components 72%, gene and proteins 64%, chemicals 53%, and biological processes and molecular functions 40%). Neji provides fast and multi-threaded data processing, annotating up to 1200 sentences/second when using dictionary-based concept identification. CONCLUSIONS: Considering the provided features and underlying characteristics, we believe that Neji is an important contribution to the biomedical community, streamlining the development of complex concept recognition solutions. Neji is freely available at http://bioinformatics.ua.pt/neji. David Campos 0001, Sérgio Matos, José Luís Oliveira |
BMC Bioinform. | 3 |
| 2013 | An innovative portal for rare genetic diseases research: The semantic DiseasecardabstractAdvances in "omics" hardware and software technologies are bringing rare diseases research back from the sidelines. Whereas in the past these disorders were seldom considered relevant, in the era of whole genome sequencing the direct connections between rare phenotypes and a reduced set of genes are of vital relevance. This increased interest in rare genetic diseases research is pushing forward investment and effort towards the creation of software in the field, and leveraging the wealth of available life sciences data. Alas, most of these tools target one or more rare diseases, are focused solely on a single type of user, or are limited to the most relevant scientific breakthroughs for a specific niche. Furthermore, despite some high quality efforts, the ever-growing number of resources, databases, services and applications is still a burden to this area. Hence, there is a clear interest in new strategies to deliver a holistic perspective over the entire rare genetic diseases research domain. This is Diseasecard's reasoning, to build a true lightweight knowledge base covering rare genetic diseases. Developed with the latest semantic web technologies, this portal delivers unified access to a comprehensive network for researchers, clinicians, patients and bioinformatics developers. With in-context access covering over 20 distinct heterogeneous resources, Diseasecard's workspace provides access to the most relevant scientific knowledge regarding a given disorder, whether through direct common identifiers or through full-text search over all connected resources. In addition to its user-oriented features, Diseasecard's semantic knowledge base is also available for direct querying, enabling everyone to include rare genetic diseases knowledge in new or existing information systems. Diseasecard is publicly available at http://bioinformatics.ua.pt/diseasecard/. Pedro Lopes 0002, José Luís Oliveira |
J. Biomed. Informatics | 2 |
| 2013 | A common API for delivering services over multi-vendor cloud resources
Luís Bastião, Carlos Costa 0001, José Luís Oliveira |
J. Syst. Softw. | 3 |
| 2012 | Dicoogle relay - a cloud communications bridge for medical imagingabstractOver the last decades, information systems for medical imaging sharing are imposing themselves as important tools for the diagnostic and study of pathologies. One of the most important advantages of those systems is to allow widespread sharing and remote access to medical data. Nevertheless, there is no simple solution for imaging data exchange between multiple places due to bureaucratic and technical issues. The paradigm introduced by Dicoogle project potentiates queries over a set of distributed repositories, which are logically indexed as a single federate unit. This paper describes a Cloud-based relay service that acts as a bridge of communications between the different institutions, allowing the community to access, share and discover imaging records. Carlos Viana-Ferreira, Carlos Costa 0001, José Luís Oliveira |
CBMS | 3 |
| 2012 | Harmonization of gene/protein annotations: towards a gold standard MEDLINEabstractMOTIVATION: The recognition of named entities (NER) is an elementary task in biomedical text mining. A number of NER solutions have been proposed in recent years, taking advantage of available annotated corpora, terminological resources and machine-learning techniques. Currently, the best performing solutions combine the outputs from selected annotation solutions measured against a single corpus. However, little effort has been spent on a systematic analysis of methods harmonizing the annotation results and measuring against a combination of Gold Standard Corpora (GSCs). RESULTS: We present Totum, a machine learning solution that harmonizes gene/protein annotations provided by heterogeneous NER solutions. It has been optimized and measured against a combination of manually curated GSCs. The performed experiments show that our approach improves the F-measure of state-of-the-art solutions by up to 10% (achieving ≈70%) in exact alignment and 22% (achieving ≈82%) in nested alignment. We demonstrate that our solution delivers reliable annotation results across the GSCs and it is an important contribution towards a homogeneous annotation of MEDLINE abstracts. AVAILABILITY AND IMPLEMENTATION: Totum is implemented in Java and its resources are available at http://bioinformatics.ua.pt/totum David Campos 0001, Sérgio Matos, Ian Lewin, José Luís Oliveira, Dietrich Rebholz-Schuhmann |
Bioinform. | 4 |
| 2012 | EuGene: maximizing synthetic gene design for heterologous expressionabstractUNLABELLED: Numerous software applications exist to deal with synthetic gene design, granting the field of heterologous expression a significant support. However, their dispersion requires the access to different tools and online services in order to complete one single project. Analyzing codon usage, calculating codon adaptation index (CAI), aligning orthologs and optimizing genes are just a few examples. A software application, EuGene, was developed for the optimization of multiple gene synthetic design algorithms. In a seamless automatic form, EuGene calculates or retrieves genome data on codon usage (relative synonymous codon usage and CAI), codon context (CPS and codon pair bias), GC content, hidden stop codons, repetitions, deleterious sites, protein primary, secondary and tertiary structures, gene orthologs, species housekeeping genes, performs alignments and identifies genes and genomes. The main function of EuGene is analyzing and redesigning gene sequences using multi-objective optimization techniques that maximize the coding features of the resulting sequence. AVAILABILITY: EuGene is freely available for non-commercial use, at http://bioinformatics.ua.pt/eugene. Paulo Gaspar, José Luís Oliveira, Jörg Frommlet, Manuel A. S. Santos, Gabriela R. Moura |
Bioinform. | 2 |
| 2012 | Automatic Filtering and Substantiation of Drug Safety SignalsabstractDrug safety issues pose serious health threats to the population and constitute a major cause of mortality worldwide. Due to the prominent implications to both public health and the pharmaceutical industry, it is of great importance to unravel the molecular mechanisms by which an adverse drug reaction can be potentially elicited. These mechanisms can be investigated by placing the pharmaco-epidemiologically detected adverse drug reaction in an information-rich context and by exploiting all currently available biomedical knowledge to substantiate it. We present a computational framework for the biological annotation of potential adverse drug reactions. First, the proposed framework investigates previous evidences on the drug-event association in the context of biomedical literature (signal filtering). Then, it seeks to provide a biological explanation (signal substantiation) by exploring mechanistic connections that might explain why a drug produces a specific adverse reaction. The mechanistic connections include the activity of the drug, related compounds and drug metabolites on protein targets, the association of protein targets to clinical events, and the annotation of proteins (both protein targets and proteins associated with clinical events) to biological pathways. Hence, the workflows for signal filtering and substantiation integrate modules for literature and database mining, in silico drug-target profiling, and analyses based on gene-disease networks and biological pathways. Application examples of these workflows carried out on selected cases of drug safety signals are discussed. The methodology and workflows presented offer a novel approach to explore the molecular mechanisms underlying adverse drug reactions. Anna Bauer-Mehren, Erik M. van Mulligen, Paul Avillach, María del Carmen Carrascosa, Ricard García-Serna, Janet Piñero González, Pedro Lopes 0002, José Luís Oliveira, Gayo Diallo, Ernst Ahlberg Helgee, Scott Boyer, Jordi Mestres, Ferran Sanz, Jan A. Kors, Laura Inés Furlong |
PLoS Comput. Biol. | 9 |
| 2012 | A RESTful Image Gateway for Multiple Medical Image RepositoriesabstractMobile technologies are increasingly important components in telemedicine systems and are becoming powerful decision support tools. Universal access to data may already be achieved by resorting to the latest generation of tablet devices and smartphones. However, the protocols employed for communicating with image repositories are not suited to exchange data with mobile devices. In this paper, we present an extensible approach to solving the problem of querying and delivering data in a format that is suitable for the bandwidth and graphic capacities of mobile devices. We describe a three-tiered component-based gateway that acts as an intermediary between medical applications and a number of Picture Archiving and Communication Systems (PACS). The interface with the gateway is accomplished using Hypertext Transfer Protocol (HTTP) requests following a Representational State Transfer (REST) methodology, which relieves developers from dealing with complex medical imaging protocols and allows the processing of data on the server side. Frederico Valente, Carlos Viana-Ferreira, Carlos Costa 0001, José Luís Oliveira |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2010 | GeneBrowser 2: an application to explore and identify common biological traits in a set of genesabstractBACKGROUND: The development of high-throughput laboratory techniques created a demand for computer-assisted result analysis tools. Many of these techniques return lists of genes whose interpretation requires finding relevant biological roles for the problem at hand. The required information is typically available in public databases, and usually, this information must be manually retrieved to complement the analysis. This process is a very time-consuming task that should be automated as much as possible. RESULTS: GeneBrowser is a web-based tool that, for a given list of genes, combines data from several public databases with visualisation and analysis methods to help identify the most relevant and common biological characteristics. The functionalities provided include the following: a central point with the most relevant biological information for each inserted gene; a list of the most related papers in PubMed and gene expression studies in ArrayExpress; and an extended approach to functional analysis applied to Gene Ontology, homologies, gene chromosomal localisation and pathways. CONCLUSIONS: GeneBrowser provides a unique entry point to several visualisation and analysis methods, providing fast and easy analysis of a set of genes. GeneBrowser fills the gap between Web portals that analyse one gene at a time and functional analysis tools that are limited in scope and usually desktop-based. Joel Arrais, João E. Pereira, José Luís Oliveira |
BMC Bioinform. | 4 |
| 2010 | Concept-based query expansion for retrieving gene related publications from MEDLINEabstractBACKGROUND: Advances in biotechnology and in high-throughput methods for gene analysis have contributed to an exponential increase in the number of scientific publications in these fields of study. While much of the data and results described in these articles are entered and annotated in the various existing biomedical databases, the scientific literature is still the major source of information. There is, therefore, a growing need for text mining and information retrieval tools to help researchers find the relevant articles for their study. To tackle this, several tools have been proposed to provide alternative solutions for specific user requests. RESULTS: This paper presents QuExT, a new PubMed-based document retrieval and prioritization tool that, from a given list of genes, searches for the most relevant results from the literature. QuExT follows a concept-oriented query expansion methodology to find documents containing concepts related to the genes in the user input, such as protein and pathway names. The retrieved documents are ranked according to user-definable weights assigned to each concept class. By changing these weights, users can modify the ranking of the results in order to focus on documents dealing with a specific concept. The method's performance was evaluated using data from the 2004 TREC genomics track, producing a mean average precision of 0.425, with an average of 4.8 and 31.3 relevant documents within the top 10 and 100 retrieved abstracts, respectively. CONCLUSIONS: QuExT implements a concept-based query expansion scheme that leverages gene-related information available on a variety of biological resources. The main advantage of the system is to give the user control over the ranking of the results by means of a simple weighting scheme. Using this approach, researchers can effortlessly explore the literature regarding a group of genes and focus on the different aspects relating to these genes. Sérgio Matos, Joel Arrais, João Maia-Rodrigues, José Luís Oliveira |
BMC Bioinform. | 4 |
| 2009 | An evaluation of network management protocolsabstractDuring the last decade several network management solutions have been proposed or extended to cope with the growing complexity of networks, systems and services. Architectures, protocols, and information models have been proposed as a way to better respond to the new and different demands of global networks. However this offer also leads to a growing complexity of management solutions and to an increase in systems' requirements. The current management landscape is populated with a multiplicity of protocols, initially developed as an answer to different requirements. This paper presents a comparative study of currently common management protocols in All-IP networks: SNMP, COPS, Diameter, CIM/XML over HTTP and CIM/XML over SOAP. This assessment was focused on wireless aspect issues, and as such includes measures of bandwidth, packets, round-trip delays, and agents' requirements. We also analyzed the advantages of compression in these protocols. Pedro Gonçalves 0001, José Luís Oliveira, Rui L. Aguiar |
Integrated Network Management | 2 |
| 2008 | Dynamic service integration using web-based workflowsabstractWeb services have been the main leverage to the development of Service Oriented Architecture (SOA), essentially a collection of interacting software agents with a loosely coupling organization. Despite several composition solutions already exist, the integration of services in a seamless and user-friendly way is not yet a complete solved problem. Most of the times, this integration is performed in hard-coded monolithic desktop applications. Pedro Lopes 0002, Joel Arrais, José Luís Oliveira |
iiWAS | 3 |
| 2003 | Critical Information Systems Authentication Based on PKC and Biometrics
Carlos Costa 0001, José Luís Oliveira, Augusto Silva |
ICWE | 2 |
| 2003 | Electronic Patient Record Virtually Unique Based on a Crypto Smart Card
Carlos Costa 0001, José Luís Oliveira, Augusto Silva |
ICWE | 2 |
| 2003 | Delegation of Expressions for Distributed SNMP Information Processing
Rui Pedro Lopes, José Luís Oliveira |
Integrated Network Management | 2 |
| 2001 | SNMP Management of MASIF PlatformsabstractIn this paper we describe the architecture, the development and the assessment results of an SNMP agent that allows MASIF (mobile agent system interoperability facilities) compliant platforms to be managed through SNMP. The major outcome of this work is a simple integration of mobile agent technology with any well-established commercial network management system. Rui Pedro Lopes, José Luís Oliveira |
Integrated Network Management | 2 |
| 1994 | A Methodology to Represent Logical Network Topology Information
José Luís Oliveira, Joaquim Arnaldo Martins |
NOMS | 1 |