Telmo de Menezes e Silva Filho

dblp:60/7361 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-0826-6885ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Language Models Do Not Embed Numbers Continuously (Student Abstract)
abstract
We evaluate how well large language model embeddings represent continuous numerical values across different precisions and ranges. Using linear models and principal component analysis on models from major providers, we show that while embeddings can reconstruct numbers with high fidelity (R2 ≥ 0.95), they introduce substantial noise, with principal components explaining less than 40% of embedding variance. Performance degrades with increasing decimal precision and mixed-sign values, revealing fundamental limitations in how these models encode numerical information.
Alex O. Davies, Roussel Desmond Nzoyem, Nirav Ajmeri, Telmo de Menezes e Silva Filho
AAAI4
2026 Topology only pre-training: towards generalised multi-domain graph models
abstract
The principal benefit of unsupervised representation learning is that a pre-trained model can be fine-tuned where data or labels are scarce. Existing approaches for graph representation learning are domain specific, maintaining consistent node and edge features across the pre-training and target datasets. This has precluded transfer to multiple domains. We present Topology Only Pre-Training (ToP), a graph pre-training method based on node and edge feature exclusion. We show positive transfer on evaluation datasets from multiple domains, including domains not present in pre-training data, running directly contrary to assumptions made in contemporary works. On 75% of experiments, ToP models perform significantly ([Formula: see text]) better than a supervised baseline. Performance is significantly positive on 85.7% of tasks when node and edge features are used in fine-tuning. We further show that out-of-domain topologies can produce more useful pre-training than in-domain. Under ToP we show better transfer from non-molecule pre-training, compared to molecule pre-training, on 79% of molecular benchmarks. Against the limited set of other generalist graph models ToP performs strongly, including against models with many orders of magnitude larger. These findings show that ToP opens broad areas of research in both transfer learning on scarcely populated graph domains and in graph foundation models. Supplementary Information: The online version contains supplementary material available at 10.1007/s10618-026-01210-1.
Alex O. Davies, Riku Green, Telmo de Menezes e Silva Filho, Nirav Ajmeri
Data Min. Knowl. Discov.3
2025 Stratify: unifying multi-step forecasting strategies
abstract
Abstract A key aspect of temporal domains is the ability to make predictions multiple time-steps into the future, a process known as multi-step forecasting (MSF). At the core of this process is selecting a forecasting strategy; however, with no existing frameworks to map out the space of strategies, practitioners are left with ad-hoc methods for strategy selection. In this work, we propose Stratify , a parameterised framework that addresses multi-step forecasting, unifying existing strategies and introducing novel, improved strategies. We evaluate Stratify on 18 benchmark datasets, five function classes, and short to long forecast horizons (10, 20, 40, 80) in the univariate setting. In over 84% of 1080 experiments, novel strategies in Stratify improved performance compared to all existing ones. Importantly, we find that no single strategy consistently outperforms others in all task settings, highlighting the need for practitioners to explore the Stratify space to carefully search and select forecasting strategies based on task-specific requirements. Our results are the most comprehensive benchmarking of known and novel forecasting strategies. We share the code ( https://github.com/zs18656/stratify_unifying_MSF ) to reproduce our results.
Riku Green, Grant Stevens, Zahraa Said Abdallah, Telmo de Menezes e Silva Filho
Data Min. Knowl. Discov.4
2025 Investigations into deep Reinforcement Learning for wind farm set-point optimisation
abstract
Wake steering is a form of wind farm flow control in which upstream turbines are deliberately yawed to misalign with the incoming wind in order to reduce the impact of wakes on downstream turbines. This technique can give a net increase in the power generated by an array of turbines compared to standard greedy control where each turbine acts for its own benefit by aligning with the incoming wind. However, optimising the set-points of multiple turbines under varying wind conditions can be prohibitively complex for traditional, white-box models. Reinforcement Learning (RL) agents learn optimal long-term behaviours through “trial-and-error”, making them suited to controlling arrays of wind turbines under changing wind conditions for maximum farm power. Related works applying RL to this problem have tended to concentrate on either single wind directions or ranges up to around ± 1 0 ∘ . Here, the Deep Deterministic Policy Gradient algorithm has been used to train RL agents to control a nine-turbine array to implement wake steering under multiple wind directions between ± 4 5 ∘ . While the agents were trained on steady-state (time-averaged) wind flow data, the performance of the final agent was tested on “quasi-dynamic” wind flow with varying wind direction. Under these conditions, the final agent achieved an average of 7% more power than greedy control per direction. This agent was then used to control the wind farm under a smaller subset of directions including many not seen during training, gaining on average 17% additional farm power per direction compared to greedy control. • Reinforcement learning agents were trained to maximise wind farm power. • Agents were trained to control the farm under a wide range of wind conditions. • A robust training configuration evolved through hyperparameter investigations. • Agent achieved on average 7% increase in farm power on conditions used in training. • Agent achieved on average 17% increase in farm power on new conditions.
Helen Sheehan, Daniel J. Poole, Telmo de Menezes e Silva Filho, Ervin Bossanyi, Lars Landberg
Expert Syst. Appl.3
2025 CLAIRE: clustering evaluation based on item response theory and model agreement
Manuel Ferreira Junior, Eufrásio de Andrade Lima Neto, Marcelo Rodrigo Portela Ferreira, Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio
Mach. Learn.4
2025 Novel applications of item response theory for analysing data set complexity and benchmark selection
João Luiz Junho Pereira, Alfredo Antonio Alencar Exposito de Queiroz, Telmo de Menezes e Silva Filho, Ana Carolina Lorena, Rafael Gomes Mantovani, Gisele L. Pappa, Ricardo B. C. Prudêncio
Mach. Learn.3
2024 Meta-Learning and Novelty Detection for Machine Learning with Reject Option
abstract
Preventing a Machine Learning (ML) predictor from making an unreliable prediction is extremely important in sensitive application domains, such as health contexts. In this sense, strategies based on the reject option have been increasingly explored. However, few studies explore the ability of meta-learning to inspect the errors of a base predictor under analysis, in such a way to generalize when the predictor is confident or not. Therefore, the current paper proposes a novel solution for ML with reject option based on the combination of meta-learning and novelty detection. The proposal addresses two distinct situations where a prediction should be rejected. First, novel detection is adopted to identify out-of-distribution instances, i.e., instances that significantly differ from those ones adopted to train the base predictor. Second, meta-learning is adopted to detect instances in regions of data where the base model has shown poor predictive performance during its evaluation. Such instances mainly lie in areas of class overlap or noisy regions in the training data. The results in experiments on synthetic and real data showed the superiority of the solution compared to those based only on meta-learning (aka without novelty detection) and those based on classifier confidence.
Patrícia Drapal, Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio
IJCNN2
2024 Assessor Models for Explaining Instance Hardness in Classification Problems
abstract
Understanding the difficulty of individual instances in a classification problem is important to define the limits of learning performance in the problem. Previous works are devoted to measuring Instance Hardness (IH), while solutions for explaining IH are still not deeply investigated. In this paper, we rely on using assessor models and eXplanaible AI (XAI) techniques to predict and explain IH. Many XAI techniques have been developed in the literature to explain the predictions of Machine Learning (ML) models. In our work, we are focused on explaining the difficulty of instances. Given a classification dataset, we trained and evaluated a pool of diverse ML models to measure the IH of each instance. Then, we trained an assessor model to predict the IH based on the instances’ features. Once the assessor is built, its predictions (i.e., the expected IH) can be explained using XAI techniques. In our experiments, we produced Partial Dependence Plots (PDP) to inspect the marginal effect of specific features on the IH predicted by the assessor. From the PDPs, we could check how IH is distributed along the instances’ features in a problem, and more specifically, we could visualize areas of high expected predictive difficulty.
Ricardo B. C. Prudêncio, Ana Carolina Lorena, Telmo de Menezes e Silva Filho, Patrícia Drapal, Maria Gabriela Valeriano
IJCNN3
2024 Recommendation systems with user and item profiles based on symbolic modal data
Delmiro D. Sampaio Neto, Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza
Neural Comput. Appl.2
2024 Assessment of volumetric dense tissue segmentation in tomosynthesis using deep virtual clinical trials
Bruno Barufaldi, Jordy V. Gomes, Telmo de Menezes e Silva Filho, Thaís Gaudencio do Rêgo, Yuri Malheiros, Trevor Lewis Vent, Aimilia Gastounioti, Andrew D. A. Maidment
Pattern Recognit.3
2023 Classifier calibration: a survey on how to assess and improve predicted class probabilities
abstract
Abstract This paper provides both an introduction to and a detailed overview of the principles and practice of classifier calibration. A well-calibrated classifier correctly quantifies the level of uncertainty or confidence associated with its instance-wise predictions. This is essential for critical applications, optimal decision making, cost-sensitive classification, and for some types of context change. Calibration research has a rich history which predates the birth of machine learning as an academic field by decades. However, a recent increase in the interest on calibration has led to new methods and the extension from binary to the multiclass setting. The space of options and issues to consider is large, and navigating it requires the right set of concepts and tools. We provide both introductory material and up-to-date technical details of the main concepts and methods, including proper scoring rules and other evaluation metrics, visualisation approaches, a comprehensive account of post-hoc calibration methods for binary and multiclass classification, and several advanced topics.
Telmo de Menezes e Silva Filho, Hao Song 0007, Miquel Perelló-Nieto, Raúl Santos-Rodríguez, Meelis Kull, Peter A. Flach
Mach. Learn.1
2023 Classifying breast lesions in Brazilian thermographic images using convolutional neural networks
Flávia R. S. Brasileiro, Delmiro D. Sampaio Neto, Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza, Marcus C. Araújo
Neural Comput. Appl.3
2022 Evaluating regression algorithms at the instance level using item response theory
João V. C. Moraes, Jessica T. S. Reinaldo, Manuel Ferreira Junior, Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio
Knowl. Based Syst.4
2022 A two-level Item Response Theory model to evaluate speech synthesis and recognition
Chaina Santos Oliveira, João V. C. Moraes, Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio
Speech Commun.3
2021 Kohonen map-wise regression applied to interval data
Leandro Carlos de Souza, Bruno A. Pimentel, Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza
Knowl. Based Syst.3
2020 Item Response Theory for Evaluating Regression Algorithms
abstract
Item Response Theory (IRT) is a tool developed in psychometrics to measure latent abilities of human respondents based on their responses to items with different levels of difficulty. Recently, IRT has been applied to evaluation in AI, by treating the algorithms as respondents and the AI tasks as items. Particularly in machine learning, IRT has been applied for evaluation of classifiers based on their predictions to each test instance. Based on a matrix of responses (classifiers vs instances), the IRT model estimates the latent difficulty and discrimination of each instance, as well as the ability of each classifier, in such a way that a classifier receives high ability value when it tends to correctly classify the most difficult instances. The IRT models previously adopted for evaluation in classification are not directly applied for regression, since they rely on dichotomous responses (i.e., a response has to be either correct or incorrect). In this paper we propose a new IRT model, particularly designed for dealing with nonnegative unbounded responses, which is adequate for modelling the absolute errors of regression algorithms. In the proposed model, responses follow a gamma distribution, parameterised according to respondents' abilities and items' difficulty and discrimination parameters. The proposed parameterisation results in item characteristic curves with more flexible shapes compared to the logistic curves widely adopted in IRT. The proposed model was evaluated with diverse regression algorithms and two benchmark datasets, one synthetic and one real. Useful insights were derived by inspecting regions in these datasets that present different levels of difficulty and discrimination.
João V. C. Moraes, Jessica T. S. Reinaldo, Ricardo B. C. Prudêncio, Telmo de Menezes e Silva Filho
IJCNN4
2019 $β^3$-IRT: A New Item Response Model and its Applications
abstract
Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the $\beta^3$-IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Curves. In experiments we applied the proposed model to data from an online exam platform, and show our model outperforms a more standard 2PL-ND model on all datasets. Furthermore, we show how to apply $\beta^3$-IRT to assess the ability of machine learning classifiers.This novel application results in a new metric for evaluating the quality of the classifier’s probability estimates, based on the inferred difficulty and discrimination of data instances.
Telmo de Menezes e Silva Filho, Ricardo B. C. Prudêncio, Tom Diethe, Peter A. Flach
AISTATS2
2019 Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with Dirichlet calibration
abstract
Class probabilities predicted by most multiclass classifiers are uncalibrated, often tending towards over-confidence. With neural networks, calibration can be improved by temperature scaling, a method to learn a single corrective multiplicative factor for inputs to the last softmax layer. On non-neural models the existing methods apply binary calibration in a pairwise or one-vs-rest fashion. We propose a natively multiclass calibration method applicable to classifiers from any model class, derived from Dirichlet distributions and generalising the beta calibration method from binary classification. It is easily implemented with neural nets since it is equivalent to log-transforming the uncalibrated probabilities, followed by one linear layer and softmax. Experiments demonstrate improved probabilistic predictions according to multiple measures (confidence-ECE, classwise-ECE, log-loss, Brier score) across a wide range of datasets and classifiers. Parameters of the learned Dirichlet calibration map provide insights to the biases in the uncalibrated model.
Meelis Kull, Miquel Perelló-Nieto, Markus Kängsepp, Telmo de Menezes e Silva Filho, Hao Song 0007, Peter A. Flach
NeurIPS4
2017 Beta calibration: a well-founded and easily implemented improvement on logistic calibration for binary classifiers
abstract
For optimal decision making under variable class distributions and misclassification costs a classifier needs to produce well-calibrated estimates of the posterior probability. Isotonic calibration is a powerful non-parametric method that is however prone to overfitting on smaller datasets; hence a parametric method based on the logistic curve is commonly used. While logistic calibration is designed for normally distributed per-class scores, we demonstrate experimentally that many classifiers including Naive Bayes and Adaboost suffer from a particular distortion where these score distributions are heavily skewed. In such cases logistic calibration can easily yield probability estimates that are worse than the original scores. Moreover, the logistic curve family does not include the identity function, and hence logistic calibration can easily uncalibrate a perfectly calibrated classifier. In this paper we solve all these problems with a richer class of calibration maps based on the beta distribution. We derive the method from first principles and show that fitting it is as easy as fitting a logistic curve. Extensive experiments show that beta calibration is superior to logistic calibration for Naive Bayes and Adaboost.
Meelis Kull, Telmo de Menezes e Silva Filho, Peter A. Flach
AISTATS2
2017 A parametrized approach for linear regression of interval data
Leandro Carlos de Souza, Renata M. C. R. de Souza, Getúlio Jose Amorim Amaral, Telmo de Menezes e Silva Filho
Knowl. Based Syst.4
2016 Background Check: A General Technique to Build More Reliable and Versatile Classifiers
abstract
We introduce a powerful technique to make classifiers more reliable and versatile. Background Check equips classifiers with the ability to assess the difference of unlabelled test data from the training data. In particular, Background Check gives classifiers the capability to (i) perform cautious classification with a reject option, (ii) identify outliers, and (iii) better assess the confidence in their predictions. We derive the method from first principles and consider four particular relationships between background and foreground distributions. One of these assumes an affine relationship with two parameters, and we show how this bivariate parameter space naturally interpolates between the above capabilities. We demonstrate the versatility of the approach by comparing it experimentally with published special-purpose solutions for outlier detection and confident classification on 41 benchmark datasets. Results show that Background Check can match and in many cases surpass the performances of specialised approaches.
Miquel Perelló-Nieto, Telmo de Menezes e Silva Filho, Meelis Kull, Peter A. Flach
ICDM2
2016 A swarm-trained k-nearest prototypes adaptive classifier with automatic feature selection for interval data
Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza, Ricardo B. C. Prudêncio
Neural Networks1
2015 Hybrid methods for fuzzy clustering based on fuzzy c-means and improved particle swarm optimization
Telmo de Menezes e Silva Filho, Bruno A. Pimentel, Renata M. C. R. de Souza, Adriano Lorena Inácio de Oliveira
Expert Syst. Appl.1
2013 Fuzzy learning vector quantization approaches for interval data
abstract
Symbolic data analysis deals with complex data types, capable of modeling internal data variability and imprecise data. This paper introduces two Fuzzy Learning Vector Quantization algorithms for interval symbolic data. One algorithm employs an interval Euclidean distance. The second uses a weighted interval Euclidean distance to try and achieve a better performance of classification when the data set is composed of classes with varying sizes, shapes and structures. The algorithms are evaluated for their performances with synthetic and real data sets. This paper aims at contributing to the area of Supervised Learning within Symbolic Data Analysis.
Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza
FUZZ-IEEE1
2012 A Weighted Learning Vector Quantization Approach for Interval Data
Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza
ICONIP (3)1
2011 Pattern classifiers with adaptive distances
abstract
This paper presents learning vector quantization classifiers with adaptive distances. The classifiers furnish discriminant class regions from the input data set that are represented by prototypes. In order to compare prototypes and patterns, the classifiers use adaptive distances that change at each iteration and are different from one class to another or from one prototype to another. Experiments with real and synthetic data sets demonstrate the usefulness of these classifiers.
Telmo de Menezes e Silva Filho, Renata M. C. R. de Souza
IJCNN1
2009 Optimized Learning Vector Quantization Classifier with an Adaptive Euclidean Distance
Renata M. C. R. de Souza, Telmo de Menezes e Silva Filho
ICANN (1)2