VLDB 2026 Research / reviewers in the wild / expert
Helena de Medeiros Caseli
dblp:15/2653
· DBLP profile ↗
15ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-3996-8599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Meta4XNLI-ptBR: Brazilian Portuguese Extension of Meta4XNLI Corpus
Karina M. Johansson, Fernanda M. Assi, Isabella da Silva, Rafael V. P. Passador, Isabela Rodrigues, Aline Paes, Helena de Medeiros Caseli |
LREC | 7 |
| 2026 | Figurative Language in Alzheimer's Discourse: Linguistic and Neural Alignment in Clinical Narratives
Diana Kylymnyk, Vitória Hilgert Tomasel, Helena de Medeiros Caseli, Edward Watkins, Aline Villavicencio, Rodrigo Wilkens |
LREC | 3 |
| 2026 | EduBench: A Portuguese Benchmark for Open-Ended Discursive Question Answering
Pedro H. Paiola, Luís Gabriel Damiati Mendes, Bruno de Oliveira Monchelato, André da Fonseca Schuck, Gabriel Lino Garcia, Douglas Rodrigues, Helena de Medeiros Caseli, João Paulo Papa |
LREC | 7 |
| 2026 | Automated Machine Learning in medical research: A systematic literature mapping study
Giovanna A. Castro, Luiza G. Barioto, Yu H. Cao, Renato Moraes Silva, Helena de Medeiros Caseli, João A. Machado-Neto, Ricardo Cerri, Aline Villavicencio, Tiago A. Almeida 0001 |
Artif. Intell. Medicine | 5 |
| 2024 | Identifying Fine-grained Depression Signs in Social Media PostsabstractNatural Language Processing has already proven to be an effective tool for helping in the identification of mental health disorders in text. However, most studies limit themselves to a binary classification setup or base their label set on pre-established resources. By doing so, they don’t explicitly model many common ways users can express their depression online, limiting our understanding of what kind of depression signs such models can accurately classify. This study evaluates how machine learning techniques deal with the classification of a fine-grained set of 21 depression signs in social media posts from Brazilian undergraduate students. We found out that model performance is not necessarily driven by a depression sign’s frequency on social media posts, since evaluated machine learning techniques struggle to classify the majority of signs of depression typically present in posts. Thus, model performance seems to be more related to the inherent difficulty of identifying a given sign than with its occurrence frequency. Augusto R. Mendes, Helena de Medeiros Caseli |
LREC/COLING | 2 |
| 2022 | Multilingual and Multimodal Learning for Brazilian PortugueseabstractHumans constantly deal with multimodal information, that is, data from different modalities, such as texts and images. In order for machines to process information similarly to humans, they must be able to process multimodal data and understand the joint relationship between these modalities. This paper describes the work performed on the VTLM (Visual Translation Language Modelling) framework from (Caglayan et al., 2021) to test its generalization ability for other language pairs and corpora. We use the multimodal and multilingual corpus How2 (Sanabria et al., 2018) in three parallel streams with aligned English-Portuguese-Visual information to investigate the effectiveness of the model for this new language pair and in more complex scenarios, where the sentence associated with each image is not a simple description of it. Our experiments on the Portuguese-English multimodal translation task using the How2 dataset demonstrate the efficacy of cross-lingual visual pretraining. We achieved a BLEU score of 51.8 and a METEOR score of 78.0 on the test set, outperforming the MMT baseline by about 14 BLEU and 14 METEOR. The good BLEU and METEOR values obtained for this new language pair, regarding the original English-German VTLM, establish the suitability of the model to other languages. Júlia Sato, Helena de Medeiros Caseli, Lucia Specia |
LREC | 2 |
| 2021 | CurL-AutoML: Curriculum Learning-based AutoMLabstractAutoML aims to find the best Machine Learning (ML) pipeline in a complex and high-dimensional search space by evaluating multiple algorithm configurations. However, training multiple ML algorithms is time-consuming, and as AutoML tools are frequently time-constrained, the exploration of the search space may find sub-optimal results. In this work, we explore the application of curriculum learning techniques to overcome this limitation. Curriculum and anti-curriculum learning have improved model performance and accelerated the training process on previous empirical investigations using optimization-based models by ordering examples during model training based on their difficulty. We apply and compare curriculum strategies on an AutoML system to accelerate the search space exploration and find good-performing machine learning pipelines efficiently. The results indicate that AutoML can benefit from a curriculum strategy. Furthermore, in most of the evaluated scenarios, the curriculum strategies led to better classification results. Lucas Nildaimon dos Santos Silva, Lucas Cardoso Silva, Fernando Rezende Zagatti, Bruno Silva Sette, Helena de Medeiros Caseli, Daniel Lucrédio, Diego Furtado Silva |
ICMLA | 5 |
| 2021 | MetaPrep: Data preparation pipelines recommendation via meta-learningabstractData preparation is a mandatory phase in the machine learning pipeline. The goal of data preparation is to convert noisy and disordered data into refined data that can be used by the algorithms. However, data preparation is time-consuming and requires specialized knowledge about the data and algorithms. Therefore, automating data preparation is essential to decrease the effort made by data scientists to develop satisfactory models. Despite its relevance, current AutoML platforms disregard or make simple hardcoded data preparation pipelines. Trying to fill this gap, we present a meta-learning-based recommendation system for data preparation. Our system recommends five pipelines, ranked by their relevance, making it useful for users with varying degrees of experience. Using the top-1 pipeline we demonstrated that our proposal allows a better performance of an AutoML system. Furthermore, the accuracy rates of our method were comparable to those achieved by a reinforcement-learning-based algorithm with the same goal, but it was up to two orders of magnitude faster. Moreover, we tested our method in a real-world application and evaluated its benefits and limitations in this scenario. Fernando Rezende Zagatti, Lucas Cardoso Silva, Lucas Nildaimon dos Santos Silva, Bruno Silva Sette, Helena de Medeiros Caseli, Daniel Lucrédio, Diego Furtado Silva |
ICMLA | 5 |
| 2021 | Committee of NAS-based modelsabstractNetwork Architecture Search (NAS) has achieved impressive results and generated models comparable with humans' classifications. Automating the definition of a neural architecture reduces the need for expert work efforts and mitigates human bias from architecture design. NAS techniques usually consist of an algorithm to search for the best architecture in a predetermined space of parameters or functions. Due to the number of deep neural architectures' parameters, this search space includes millions of parameters, which makes NAS a cost procedure and may lead the search to overfit the training set. To reduce NAS search spaces' complexity and still obtain competitive results, we propose CoNAS, a committee of NAS-based models, by restricting the search spaces to perform Differentiable ARchiTecture Search (DARTS). Our results point to improved accuracy over DARTS on CIFAR-10, training the networks from scratch. and Imagnette, using a transfer learning approach. Bruno Silva Sette, Lucas Cardoso Silva, Fernando Rezende Zagatti, Lucas Nildaimon dos Santos Silva, Daniel Lucrédio, Helena de Medeiros Caseli, Diego Furtado Silva |
IJCNN | 6 |
| 2020 | Benchmarking Machine Learning Solutions in ProductionabstractMachine learning (ML) is becoming critical to many businesses. Keeping an ML solution online and responding is therefore a necessity, and is part of the MLOps (Machine Learning operationalization) movement. One aspect for this process is monitoring not only prediction quality, but also system resources. This is important to correctly provide the necessary infrastructure, either using a fully-managed cloud platform or a local solution. This is not a difficult task, as there are many tools available. However, it requires some planning and knowledge about what to monitor. Also, many ML professionals are not experts in system operations and may not have the skills to easily setup a monitoring and benchmarking environment. In the spirit of MLOps, this paper presents an approach, based on a simple API and set of tools, to monitor ML solutions. The approach was tested with 9 different solutions. The results indicate that the approach can deliver useful information to help in decision making, proper resource provision and operation of ML systems. Lucas Cardoso Silva, Fernando Rezende Zagatti, Bruno Silva Sette, Lucas Nildaimon dos Santos Silva, Daniel Lucrédio, Diego Furtado Silva, Helena de Medeiros Caseli |
ICMLA | 7 |
| 2020 | NMT and PBSMT Error Analyses in English to Brazilian Portuguese Automatic TranslationsabstractMachine Translation (MT) is one of the most important natural language processing applications. Independently of the applied MT approach, a MT system automatically generates an equivalent version (in some target language) of an input sentence (in some source language). Recently, a new MT approach has been proposed: neural machine translation (NMT). NMT systems have already outperformed traditional phrase-based statistical machine translation (PBSMT) systems for some pairs of languages. However, any MT approach outputs errors. In this work we present a comparative study of MT errors generated by a NMT system and a PBSMT system trained on the same English – Brazilian Portuguese parallel corpus. This is the first study of this kind involving NMT for Brazilian Portuguese. Furthermore, the analyses and conclusions presented here point out the specific problems of NMT outputs in relation to PBSMT ones and also give lots of insights into how to implement automatic post-editing for a NMT system. Finally, the corpora annotated with MT errors generated by both PBSMT and NMT systems are also available. Helena de Medeiros Caseli, Marcio Lima Inácio |
LREC | 1 |
| 2018 | The Effects of Unimodal Representation Choices on Multimodal Learning
Fernando Tadao Ito, Helena de Medeiros Caseli, Jander Moreira |
LREC | 2 |
| 2015 | Automatic machine translation error identification
Débora Beatriz de Jesus Martins, Helena de Medeiros Caseli |
Mach. Transl. | 2 |
| 2014 | Automatic semantic relation extraction from Portuguese texts
Leonardo Sameshima Taba, Helena de Medeiros Caseli |
LREC | 2 |
| 2006 | Automatic induction of bilingual resources from aligned parallel corpora: application to shallow-transfer machine translation
Helena de Medeiros Caseli, Maria das Graças Volpe Nunes, Mikel L. Forcada |
Mach. Transl. | 1 |