Fernando Bação

dblp:52/5976 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-0834-0275ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Machine learning methods for detecting smart contracts vulnerabilities within Ethereum blockchain - A review
abstract
This paper presents a comprehensive exploration of the intersection between machine learning and smart contract vulnerabilities on the Ethereum blockchain. Introduced by Vitalik Buterin in 2015, Ethereum stands as a prominent blockchain network, necessitating innovative approaches to secure smart contracts against vulnerabilities and potential attacks. This research follows PRISMA guidelines, posing three fundamental questions and conducting a meticulous literature review. The study categorises machine learning applications into seven distinct groups, analysing their taxonomy, feature types, and engineering methods. The findings indicate a dynamic landscape characterised by a noticeable trend towards increased complexity. This complexity is evident not only in the integration of machine learning frameworks that combine different architectures of deep learning models, such as Convolutional Neural Networks (CNN), Graph Neural Networks (GNN), or Recurrent Neural Networks (RNN), but also in the incorporation of various types of data related to smart contracts (SCs). The discussion dissects the advantages, limitations, and future directions in securing smart contracts using machine learning. The paper concludes by emphasising the evolving role of machine learning in strengthening the Ethereum blockchain, fostering trust, and enhancing security in decentralised systems.
João Crisóstomo, Fernando Bação, Victor Sousa Lobo
Expert Syst. Appl.2
2024 WSMOTER: a novel approach for imbalanced regression
abstract
Abstract Although the imbalanced learning problem is best known in the context of classification tasks, it also affects other areas of learning algorithms, such as regression. For regression, the problem is characterized by the existence of a continuous target variable domain and the need for models capable of making accurate predictions about rare events. Furthermore, such rare events with a real-value target are often the ones with greater interest in having models that can predict them. In this paper, we propose the novel approach WSMOTER (Weighting SMOTE for Regression) to tackle the imbalanced regression problem, which, according to the experimental work we present, outperforms currently available solutions to the problem.
Luís Camacho, Fernando Bação
Appl. Intell.2
2024 Correction to: WSMOTER: a novel approach for imbalanced regression
abstract
Camacho, L., & Bacao, F. (2024). Correction to: WSMOTER: a novel approach for imbalanced regression. Applied Intelligence, 54, 11160. https://doi.org/10.1007/s10489-024-05704-7
Luís Camacho, Fernando Bação
Appl. Intell.2
2024 A cross-chain access control mechanism based on blockchain and the threshold Paillier cryptosystem
abstract
With the continuous maturation of blockchain technology and the increasing demands for various industry applications, data sharing and interoperability among different blockchain networks face significant challenges. Research on cross-chain interoperability mechanisms has facilitated data collaboration across organizations and industries, enhancing the value and utility of data. When engaging in cross-chain data interactions, access control ensures the security and privacy of data while promoting collaboration and information exchange among multiple chains. Attribute-based access control can provide fine-grained authorization support, matching complex business scenarios. However, publicly disclosed policies and attributes in a transparent blockchain network may pose privacy and security issues. To address these issues, this paper proposes a cross-heterogeneous multichain data access control scheme based on attributes and threshold homomorphic encryption, achieving fine-grained and secure cross-domain access control in cross-chain networks. This scheme uses the threshold Paillier cryptosystem to encrypt and conceal user attributes and access policies. Through smart contracts, homomorphic differential computation is performed on the ciphertext of policies and attributes, to protect data privacy. This solution leverages private key decryption shares from multiple relay nodes in the cross-chain network to jointly decrypt the computation results, providing secure access control in complex cross-chain scenarios. Security analysis and experimental results demonstrate that the proposed scheme ensures security with reasonable computational overhead, to meet the access control requirements in cross-chain networks.
Haiping Si, Weixia Li, Chuanhu Zhang, Fernando Bação, Changxia Sun
Comput. Commun.7
2024 UMAP-SMOTENC: A simple, efficient, and consistent alternative for privacy-aware synthetic data generation
abstract
The intensification of governmental legislation and the social awareness around data privacy protection severely constrains organizations' data utilization capabilities. As a result, the interest in data anonymization techniques, which should preserve the patterns present in the original data but mitigate the risks of privacy leakage, has also increased. While conventional methods may compromise privacy, recently proposed deep learning generative approaches are computationally expensive and unreliable when used in tabular datasets, hindering the democratization and usability of data. In this paper, we explore this trade-off between privacy and the quality of the anonymized data, establishing a new equilibrium obtained using a synthetic oversampling technique, SMOTE-NC, on a non-linear compressed version of the input space, achieved with the application of UMAP. The introduced approach, UMAP-SMOTENC, constitutes an efficient and consistent solution that can be used without significant efforts on hyperparameter tuning or resourcing to massive computing infrastructures. An experiment was conducted to evaluate the robustness of the proposed solution, comparing several metrics and models across eight datasets with diverse characteristics. The results achieved suggest that the presented method can efficiently synthesize privacy-aware data while conserving the relevant patterns of the real dataset, particularly those required for classification tasks.
Goncalo Almeida, Fernando Bação
Knowl. Based Syst.2
2023 MapIntel: A visual analytics platform for competitive intelligence
abstract
Abstract Competitive Intelligence allows an organization to keep up with market trends and foresee business opportunities. This practice is mainly performed by analysts scanning for any piece of valuable information in a myriad of dispersed and unstructured sources. Here we present MapIntel, a system for acquiring intelligence from vast collections of text data by representing each document as a multidimensional vector that captures its own semantics. The system is designed to handle complex Natural Language queries and visual exploration of the corpus, potentially aiding overburdened analysts in finding meaningful insights to help decision‐making. The system searching module uses a retriever and re‐ranker engine that first finds the closest neighbours to the query embedding and then sifts the results through a cross‐encoder model that identifies the most relevant documents. The browsing or visualization module also leverages the embeddings by projecting them onto two dimensions while preserving the multidimensional landscape, resulting in a map where semantically related documents form topical clusters which we capture using topic modelling. This map aims at promoting a fast overview of the corpus while allowing a more detailed exploration and interactive information encountering process. We evaluate the system and its components on the 20 newsgroups data set, using the semantic document labels provided, and demonstrate the superiority of Transformer‐based components. Finally, we present a prototype of the system in Python and show how some of its features can be used to acquire intelligence from a news article corpus we collected during a period of 8 months.
Fernando Bação
Expert Syst. J. Knowl. Eng.2
2023 Geometric SMOTE for imbalanced datasets with nominal and continuous features
abstract
Imbalanced learning can be addressed in 3 different ways: Resampling, algorithmic modifications and cost-sensitive solutions. Resampling, and specifically oversampling, are more general approaches when opposed to algorithmic and cost-sensitive methods. Since the proposal of the Synthetic Minority Oversampling TEchnique (SMOTE), various SMOTE variants and neural network-based oversampling methods have been developed. However, the options to oversample datasets with nominal and continuous features are limited. We propose Geometric SMOTE for Nominal and Continuous features (G-SMOTENC), based on a combination of G-SMOTE and SMOTENC. Our method modifies SMOTENC’s encoding and generation mechanism for nominal features while using G-SMOTE’s data selection mechanism to determine the center observation and k-nearest neighbors and generation mechanism for continuous features. G-SMOTENC’s performance is compared against SMOTENC’s along with two other baseline methods, a State-of-the-art oversampling method and no oversampling. The experiment was performed over 20 datasets with varying imbalance ratios, number of metric and non-metric features and target classes. We found a significant improvement in classification performance when using G-SMOTENC as the oversampling method. An open-source implementation of G-SMOTENC is made available in the Python programming language.
João Fonseca, Fernando Bação
Expert Syst. Appl.2
2023 Improving Active Learning Performance through the Use of Data Augmentation
abstract
Active learning (AL) is a well‐known technique to optimize data usage in training, through the interactive selection of unlabeled observations, out of a large pool of unlabeled data, to be labeled by a supervisor. Its focus is to find the unlabeled observations that, once labeled, will maximize the informativeness of the training dataset, therefore reducing data‐related costs. The literature describes several methods to improve the effectiveness of this process. Nonetheless, there is a paucity of research developed around the application of artificial data sources in AL, especially outside image classification or NLP. This paper proposes a new AL framework, which relies on the effective use of artificial data. It may be used with any classifier, generation mechanism, and data type and can be integrated with multiple other state‐of‐the‐art AL contributions. This combination is expected to increase the ML classifier’s performance and reduce both the supervisor’s involvement and the amount of required labeled data at the expense of a marginal increase in computational time. The proposed method introduces a hyperparameter optimization component to improve the generation of artificial instances during the AL process as well as an uncertainty‐based data generation mechanism. We compare the proposed method to the standard framework and an oversampling‐based active learning method for more informed data generation in an AL context. The models’ performance was tested using four different classifiers, two AL‐specific performance metrics, and three classification performance metrics over 15 different datasets. We demonstrated that the proposed framework, using data augmentation, significantly improved the performance of AL, both in terms of classification performance and data selection efficiency (all the codes and preprocessed data developed for this study are available at https://github.com/joaopfonseca/publications/ ).
João Fonseca, Fernando Bação
Int. J. Intell. Syst.2
2023 Automation of Legal Precedents Retrieval: Findings from a Literature Review
abstract
Judges frequently rely their reasoning on precedents. Courts must preserve uniformity in decisions while, depending on the legal system, previous cases compel rulings. The search for methods to accurately identify similar previous cases is not new and has been a vital input, for example, to case‐based reasoning (CBR) methodologies. This literature review offers a comprehensive analysis of the advancements in automating the identification of legal precedents, primarily focusing on the paradigm shift from manual knowledge engineering to the incorporation of Artificial Intelligence (AI) technologies such as natural language processing (NLP) and machine learning (ML). While multiple approaches harnessing NLP and ML show promise, none has emerged as definitively superior, and further validation through statistically significant samples and expert‐provided ground truth is imperative. Additionally, this review employs text‐mining techniques to streamline the survey process, providing an accurate and holistic view of the current research landscape. By delineating extant research gaps and suggesting avenues for future exploration, this review serves as both a summation and a call for more targeted, empirical investigations.
Hugo Silva 0007, Nuno Antonio, Fernando Bação
Int. J. Intell. Syst.3
2022 Geometric SMOTE for regression
Luís Camacho, Georgios Douzas, Fernando Bação
Expert Syst. Appl.3
2021 G-SOMO: An oversampling approach based on self-organized maps and geometric SMOTE
Georgios Douzas, Rene Rauch, Fernando Bação
Expert Syst. Appl.3
2019 Gamification: A key determinant of massive open online course (MOOC) success
Manuela Aparicio, Tiago Oliveira 0001, Fernando Bação, Marco Painho
Inf. Manag.3
2019 Geometric SMOTE a geometrically enhanced drop-in replacement for SMOTE
Georgios Douzas, Fernando Bação
Inf. Sci.2
2018 Effective data generation for imbalanced learning using conditional generative adversarial networks
Georgios Douzas, Fernando Bação
Expert Syst. Appl.2
2018 Improving imbalanced learning through a heuristic oversampling method based on k-means and SMOTE
Georgios Douzas, Fernando Bação, Felix Last
Inf. Sci.2
2018 The Global Digital Divide: Evidence and Drivers
abstract
This article presents an analysis of the global digital divide, based on data collected from 45 countries, including the ones belonging to the European Union, OECD, Brazil, Russia, India, and China (BRIC). The analysis shows that one factor can explain a large part of the variation in the seven ICT variables used to measure the digital development of countries. This measure is then used with additional variables, which are hypothesised as drivers of the divide for a regression analysis using data from 2015, 2013, and 2011, which reveals economic and educational imbalances between countries, along with some aspects of geography, as drivers of the digital divide. Contrary to the authors' expectations, the English language is not a driver.
Frederico Cruz-Jesus, Tiago Oliveira 0001, Fernando Bação
J. Glob. Inf. Manag.3
2017 Self-Organizing Map Oversampling (SOMO) for imbalanced data set learning
Georgios Douzas, Fernando Bação
Expert Syst. Appl.2
2012 Digital divide across the European Union
Frederico Cruz-Jesus, Tiago Oliveira 0001, Fernando Bação
Inf. Manag.3
2009 GeoSOM Suite: A Tool for Spatial Clustering
Roberto Henriques, Fernando Bação, Victor Sousa Lobo
ICCSA (1)2
2009 Carto-SOM: cartogram creation using self-organizing maps
Roberto Henriques, Fernando Bação, Victor Sousa Lobo
Int. J. Geogr. Inf. Sci.2
2005 Applying genetic algorithms to zone design
Fernando Bação, Victor Sousa Lobo, Marco Painho
Soft Comput.1
2005 Exploring spatial data through computational intelligence: a joint perspective
Marco Painho, Athanasios V. Vasilakos, Fernando Bação, Witold Pedrycz
Soft Comput.3