Matilde Pato

dblp:299/5002 · also M. P. M. Pato, Matilde Pós-de-Mina Pato · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0001-8976-7651ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 9 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Why do women pursue a Ph.D. in Computer Science?
abstract
Context: Computer science, even now, attracts a small number of women, and the proportion of women in the field decreases through advancing career stages. Consequently, few women progress to Ph.D. studies in computer science after completing master’s studies. Empowering women at this stage in their careers is essential, not just for equality reasons, but to unlock untapped potential for society, industry and academia. Objective: This paper aims to identify students’ career assumptions and information related to Ph.D. studies focused on gender-based differences. We propose a program to inform female master students about Ph.D. studies that explains the process, clarifies misconceptions, and alleviates concerns. Method: An extensive survey was conducted to identify factors that encourage and discourage students from undertaking Ph.D. studies. The analysis identified statistically significant differences between those who undertook Ph.D. studies and those who did not, as well as statistically significant gender differences. A catalogue of questions to initiate discussions with potential Ph.D. students which allowed them to explore these factors was developed. These were structured into a Women’s Career Lunch program where students can explore and discuss the benefits of Ph.D. study. Results: Encouraging factors towards Ph.D. study include interest and confidence in research arising from a research involvement during earlier studies; enthusiasm for and self-confidence in computer science in addition to an interest in an academic career; encouragement from external sources; and a positive perception towards Ph.D. studies which can involve achieving personal goals. Discouraging factors include uncertainty and lack of knowledge of the Ph.D. process, a perception of lower job flexibility, and the requirement for long-term commitment. Gender differences highlighted that female students who pursue a Ph.D. have less confidence in their technical skills than males but a higher preference for interdisciplinary areas. Female students are less inclined than males to perceive the industry as offering better job opportunities and more flexible career paths than academia. Conclusions: The insights collected from the survey facilitated the development of a questions catalogue structured into the Women Career Lunch program to help students make a more informed decision concerning whether they should pursue a Ph.D. in computer science. Localised versions of this program, in 8 languages, were created to support its adoption in different countries and assist in mitigating the female under-representation challenge.
Erika Ábrahám, Miguel Goulão, Milena Vujosevic-Janicic, Sarah Jane Delany, Amal Mersni, Oleksandra Yeremenko, Ozge Buyukdagli, Karima Boudaoud, Caroline Oehlhorn, Ute Schmid, Christina Büsing, Helen Bolke-Hermanns, Kaja Köhnle, Matilde Pato, Deniz Sunar Cerci, Larissa Schmid
J. Syst. Softw.14
2025 Mapping Drug Interactions and Therapeutic Clusters through Knowledge Graph Visualization
abstract
Knowledge graphs (KGs) have emerged as powerful tools for biomedical research, enabling the integration and analysis of heterogeneous data sources. This study explores how a KG, combined with the Neo4j graph database, supports drug analysis and the extraction of biomedical insights. First, data from DailyMed, Purple Book and Orange Book are collected and standardized to ensure interoperability. Second, Named Entity Recognition is applied to address inconsistencies across sources. A hierarchical approach is used to link drugs to Disease Ontology, Orphanet, DrugBank, and ChEBI or active ingredient-based IDs, ensuring data accuracy. The constructed KG facilitates diverse analytical tasks, including the identification of drug-disease associations, longitudinal analysis of drug approval trends, and characterization of common routes of administration. Our results reveal complex interconnections between 561 drugs and 176 diseases, identifying significant regulatory hubs and therapeutic clusters. Temporal analysis demonstrated an acceleration in regulatory activity in recent years, while bipartite network analysis of administration routes revealed predominant delivery methods. Furthermore, graph algorithms provided by Neo4j allow advanced analyses such as finding the shortest paths between drugs based on their regulatory and therapeutic properties, revealing clusters of similar medications, and uncovering candidates for drug repurposing. The findings highlight the potential of KG methodologies in pharmaceutical research, offering a scalable approach to complex biomedical data analysis.
Ana Carolina Pereira, Matilde Pato, Nuno Datia
IV2
2024 Explainable Feature Ranking Using Interactive Dashboards
abstract
In the dynamic realm of machine learning, achieving transparency and understandability is crucial for fostering trust and facilitating broader adoption. This study presents an enhanced version of the Ensemble Feature Ranking algorithm, tailored to optimize feature selection in machine learning models. This paper proposes the use of an interactive dashboard application, as part of learning environment, designed to provide users with a visually intuitive platform for exploring the algorithm's internal metrics and rankings. The dashboard facilitates a deeper understanding of feature importance and algorithm behaviour, bridging the gap between complex algorithms and user comprehension. By combining advanced algorithmic techniques with a user-centric interface, our approach promotes transparency, accountability and increased user engagement in the explanation of machine learning models.
Diogo Amorim, Matilde Pato, Nuno Datia
IV2
2024 Understanding Portuguese Users of Parcel Locker Services
abstract
The rise of e-commerce, greatly accelerated by the COVID-19 pandemic, has created a need for more efficient and dependable delivery methods. As a result, alternative delivery points, such as parcel lockers, are being explored as effective solutions for distributing e-commerce products. Examining the social context, distribution, and density of parcel lockers is crucial in highlighting their importance. Previous research has identified the factors that influence the selection of delivery locations, assessed environmental risks and compared delivery methods. This study confidently analyses locker usage patterns, by studying 750 lockers and more than 200,000 parcels, as well as load and turnover rates, by delving into the demographic characteristics of Portuguese parishes, focusing on age, education, and employment status. Real-world data from a prominent Portuguese parcel locker provider reveals that education and employment status significantly impact the selection of parcel locker locations.
Matilde Pato, António Serrador, Rogério Campos-Rebelo, Nuno Datia, José Simão, Pedro Sampaio 0002
IV2
2024 Enhancing Drug Reviews Insights through Exploratory Data Analysis and Sentiment Analysis
abstract
The increasing volume of user-generated content across various online platforms has created vast datasets in multiple domains, including healthcare. This article explores the significant roles of data visualisation and sentiment analysis within the healthcare sector using the UCI ML Drug Review dataset. Our study highlights the value of exploratory data analysis and sentiment analysis in comprehending patient feedback, enriching insights from the dataset. Data visualisation effectively elucidates the data's distribution and key characteristics, while sentiment analysis, performed using TextBlob and VADER, categorises the emotional tone of patient reviews. Our methodology aims to provide a deeper understanding of patient satisfaction and medication efficacy based on user-generated content.
Ana Sofia Pinto, Matilde Pato, Nuno Datia
IV2
2023 Data Visualisation on a Mobile App for Real-Time Mental Health Monitoring
abstract
Anxiety disorders refer to mental health conditions characterized by excessive and persistent worry, fear or dread, which can interfere with daily life. These disorders are pervasive worldwide but can be treated with psychotherapy, medication, or both. Detecting them in time is key to avoid further severe development after an initial crisis. In this paper, we present a solution based on self-monitoring, using wearables, smartphones and Machine Learning (ML) to assess users' anxiety and panic levels. This is the first paper to embrace and present a visualisation system for mental health focused on patients. User studies support the solution's quality. With this system, patients are empowered to control their situation better, helping the medical staff to get more insight into when and where a crisis occurs.
Nuno Gomes, Matilde Pato, André Lourenço, Renato Marcelo, Nuno Datia
IV2
2023 NLP for Enterprise Asset Management: An Emerging Paradigm
abstract
In the field of asset management, a Work Order refers to a document that outlines the necessary steps to carry out a maintenance operation on a specific physical asset. The text on this Work orders providing details about the problem and the actions required are open-ended, not normalized, and Technician’ dependant, presenting challenges for automating asset management Work Order processing. To address the issue of automating the analysis of Work Orders, Natural Language Processing techniques are employed to process the content of these documents. The aim is to identify and extract relevant information related to actions and components within the sentences. This paper presents the Reliability Centred Maintenance for Assets solution, which utilizes a semi-automatic, human-in-the-loop approach to determine a standardised and condensed set of actions and components. The results indicate a significant increase in the number of annotations, reaching a ratio of 1:14. By implementing this solution, the manual workload associated with analysing Work Orders can be reduced, thereby improving decision support and analytical processing of the data contained within these documents.
Nuno Datia, Matilde Pato, José S. Sobral Neto, Nuno Gomes, Noel Leitão, Manuel R. Ferreira
IV3
2022 Comparing Word Embeddings through Visualisation
abstract
Asset management is a branch of facilities management that is responsible for the operation and maintenance of assets. The most common means of managing assets and their life-cycle is through requests and work orders. A request is used to report an occurrence that is detected either by a sensory device, a technician, or non-technical personnel; they are used to pointing out that something is wrong in a given asset, and needs appropriate attention. Depending on the problem, a request can give rise to a work order if the solution is not trivial. Work orders consist in technical reports that specify the asset that needs intervention and has the details about the work to be done or, in the case that the work is unknown from the start, the characteristics of the malfunctioning. Work orders contain a set of words, free text, that are not restricted from a fixed set of vocabulary, making it difficult to automatically analyse them. In this paper, we discuss the application of modern Natural Language Processing techniques to process the work order's description, while presenting a comparison between two Word Embedding models - Word2Vec and Fasttext- through semantic similarity tests between the encoded words, and a visualisation of the vector space through dimensionality reduction of the encoded vectors. The results show a better performance of the Fasttext approach, considering the semantics of the results.
Nuno Datia, Matilde Pato, José S. Sobral Neto
IV3
2022 Traffic Flow Indicator: Predicting Jams in a City
abstract
Road traffic inside cities is responsible for noise and pollution, that causes health problems, fuel consumption and waste of time in jams. Mitigation solutions are usually used to soften the impact of this problem in most cities. In particular, the city of Lisbon has taken measures to reduce pollution by closing areas of the city to the most polluting cars - the zero emission zones. However, the city still lacks visual analytics support for traffic decisions in real-time. In this paper we present a traffic flow indicator that can indicate the road traffic fluidity inside a region of interest for a given time frame, and integrated it into a interactive dashboard supported by a predictive model. With this solution, decision makers can analyse historical data and predict short-term traffic behaviour.
João Vaz, Nuno Datia, Matilde Pato, João Moura Pires
IV3
2021 Creating Recommender Systems Datasets in Scientific Fields
abstract
Recommender systems (RS) have been successfully explored in a vast number of domains, e.g. movies and tv shows, music, or e-commerce. In these domains we have a large number of datasets freely available for testing and evaluating new recommender algorithms. For example, Movielens and Netflix datasets for movies, Spotify for music, and Amazon for e-commerce, which translates into a large number of algorithms applied to these fields. In scientific fields, such as Health and Chemistry, standard and open access datasets with the information about the preferences of the users are scarce. First, it is important to understand the application domain, i.e. "what the recommended item is". Second, who are the end users: researchers, pharmacists, clinicians or policy makers. Third, the availability of data. Thus, if we wish to develop an algorithm for recommending scientific items, we do not have access to datasets with information about the past preferences of a group of users. Given this limitation, we developed a methodology, called LIBRETTI - LIterature Based RecommEndaTion of scienTific Items, whose goal is the creation of datasets, related with scientific fields. These datasets are created based on the major resource of knowledge that Science has: scientific literature. We consider the users as the authors of the publications, the items as the scientific entities (for example chemical compounds or diseases), and the ratings as the number of publications an author wrote about an entity. In this tutorial we will approach state-of-the-art recommender systems in scientific fields, explain what is Named Entity Recognition/Linking (NER/NEL) in research literature, and to demonstrate how to create a dataset for recommending drugs and diseases through research literature related to COVID-19. Our goal is to spread the use of LIBRETTI methodology in order to help in the development of recommender algorithms in scientific fields. These datasets are created based on the major resource of knowledge that Science has: scientific literature. We consider the users as the authors of the publications, the items as the scientific entities (for example chemical compounds or diseases), and the ratings as the number of publications an author wrote about an entity. In this tutorial we will approach state-of-the-art recommender systems in scientific fields, explain what is Named Entity Recognition/Linking (NER/NEL) in research literature, and to demonstrate how to create a dataset for recommending drugs and diseases through research literature related to COVID-19. Our goal is to spread the use of LIBRETTI methodology in order to help in the development of recommender algorithms in scientific fields. More info about the tutorial at https://lasigebiotm.github.io/RecSys.Scifi/.
Márcia Barros, Francisco M. Couto, Matilde Pato, Pedro Ruas
KDD3
2020 Exploring air quality using a multiple spatial resolution dashboard - a case study in Lisbon
abstract
Air quality is monitored using data recollected using fixed selected stations in a region, generally a city. Such approach does not support a fine-grained comprehension about the air quality, namely, in areas distant from the collector's stations, specially in residential urban places. In this paper, we describe a platform that will provide to city council decision-makers a visualization of air pollution data, using an interactive map-based dashboard with multiple spatial resolution. The air quality data is collected using low-cost portable sensors. Air pollution data is then integrated with other environmental contextual data and displayed into the dashboard. Such data includes, among other, spatio-temporal mobility data, providing contextual information about air pollution. The solution is tailored to city council decision-makers enabling a better understanding of air quality issues, and acting as a supporting tool for different communities, exploiting synergies to promote the sustainability of the city.
Ruben Taborda, Nuno Datia, Matilde Pato, João Moura Pires
IV3