Thiago Castro Ferreira

dblp:144/6801 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0003-0200-3646ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scaling Up Data-to-Text Generation to Longer Sequences: A New Dataset and Benchmark Results for Generation from Large Triple Sets
abstract
The ability of LLMs to write coherent, faithful long texts from structured data inputs remains relatively uncharted, in part because nearly all public data-to-text datasets contain only short input-output pairs. To address these gaps, we benchmark six LLMs, a rule‐based system and human-written texts on a new long-input dataset in English and Irish via LLM-based evaluation. We find substantial differences between models and languages.
Chinonso Cynthia Osuji, Simon Mille, Ornait O'Connell, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001
INLG4
2025 Are Multi-Agents the new Pipeline Architecture for Data-to-Text Systems?
abstract
Large Language Models (LLMs) have achieved remarkable results in natural language generation, yet challenges remain in data-to-text (D2T) tasks, particularly in controlling output, ensuring transparency, and maintaining factual consistency with the input. We introduce the first LLM-based multi-agent framework for D2T generation, coordinating specialized agents to produce high-quality, interpretable outputs. Our system combines the reasoning and acting abilities of ReAct agents, the self-correction of Reflexion agents, and the quality assurance of Guardrail agents, all directed by an Orchestrator agent that assigns tasks to three specialists—content ordering, text structuring, and surface realization—and iteratively refines outputs based on Guardrail feedback. This closed-loop design enables precise control and dynamic optimization, yielding text that is coherent, accurate, and grounded in the input data. On a relatively simple dataset like WebNLG, our framework performs competitively with end-to-end systems, highlighting its promise for more complex D2T scenarios.
Chinonso Cynthia Osuji, Brian Timoney, Mark Andrade, Thiago Castro Ferreira, Brian Davis 0001
INLG4
2024 A Persona-Based Corpus in the Diabetes Self-Care Domain - Applying a Human-Centered Approach to a Low-Resource Context
abstract
While Natural Language Processing (NLP) models have gained substantial attention, only in recent years has research opened new paths for tackling Human-Computer Design (HCD) from the perspective of natural language. We focus on developing a human-centered corpus, more specifically, a persona-based corpus in a particular healthcare domain (diabetes mellitus self-care). In order to follow an HCD approach, we created personas to model interpersonal interaction (expert and non-expert users) in that specific domain. We show that an HCD approach benefits language generation from different perspectives, from machines to humans - contributing with new directions for low-resource contexts (languages other than English and sensitive domains) where the need to promote effective communication is essential.
Rossana Cunha, Thiago Castro Ferreira, Adriana S. Pagano, Fábio Alves
LREC/COLING2
2024 Pipeline Neural Data-to-text with Large Language Models
abstract
Previous studies have highlighted the advantages of pipeline neural architectures over endto-end models, particularly in reducing text hallucination.In this study, we extend prior research by integrating pretrained language models (PLMs) into a pipeline framework, using both fine-tuning and prompting methods.Our findings show that fine-tuned PLMs consistently generate high quality text, especially within end-to-end architectures and at intermediate stages of the pipeline across various domains.These models also outperform promptbased ones on automatic evaluation metrics but lag in human evaluations.Compared to the standard five-stage pipeline architecture, a streamlined three-stage pipeline, which only include ordering, structuring, and surface realization, achieves superior performance in fluency and semantic adequacy according to the human evaluation.
Chinonso Cynthia Osuji, Brian Timoney, Thiago Castro Ferreira, Brian Davis 0001
INLG3
2024 aiXplain SDK: A High-Level and Standardized Toolkit for AI Assets
abstract
The aiXplain SDK 1 is an open-source Python toolkit which aims to simplify the wide and complex ecosystem of AI resources.The toolkit enables access to a wide selection of AI assets, including datasets, models, and metrics, from both academic and commercial sources, which can be selected, executed and evaluated in one place through different services in a standardized format with consistent documentation provided.The study showcases the potential of the proposed toolkit with different code examples and by using it on a user journey where state-of-the-art Large Language Models are fine-tuned on instruction prompt datasets, outperforming their base versions.
Shreyas Sharma, Lucas Pavanelli, Thiago Castro Ferreira, Mohamed Al-Badrashiny, Hassan Sawaf
INLG3
2023 NoRefER: a Referenceless Quality Metric for Automatic Speech Recognition via Semi-Supervised Language Model Fine-Tuning with Contrastive Learning
Kamer Ali Yüksel, Thiago Castro Ferreira, Golara Javadi, Mohamed Al-Badrashiny, Ahmet Gunduz
INTERSPEECH2
2023 Neural Data-to-Text Generation Based on Small Datasets: Comparing the Added Value of Two Semi-Supervised Learning Approaches on Top of a Large Language Model
abstract
Abstract This study discusses the effect of semi-supervised learning in combination with pretrained language models for data-to-text generation. It is not known whether semi-supervised learning is still helpful when a large-scale language model is also supplemented. This study aims to answer this question by comparing a data-to-text system only supplemented with a language model, to two data-to-text systems that are additionally enriched by a data augmentation or a pseudo-labeling semi-supervised learning approach. Results show that semi-supervised learning results in higher scores on diversity metrics. In terms of output quality, extending the training set of a data-to-text system with a language model using the pseudo-labeling approach did increase text quality scores, but the data augmentation approach yielded similar scores to the system without training set extension. These results indicate that semi-supervised learning approaches can bolster output quality and diversity, even when a language model is also present.
Chris van der Lee, Thiago Castro Ferreira, Chris Emmery, Travis J. Wiltshire, Emiel Krahmer
Comput. Linguistics2
2022 Generating Questions from Wikidata Triples
abstract
Question generation from knowledge bases (or knowledge base question generation, KBQG) is the task of generating questions from structured database information, typically in the form of triples representing facts. To handle rare entities and generalize to unseen properties, previous work on KBQG resorted to extensive, often ad-hoc pre- and post-processing of the input triple. We revisit KBQG – using pre training, a new (triple, question) dataset and taking question type into account – and show that our approach outperforms previous work both in a standard and in a zero-shot setting. We also show that the extended KBQG dataset (also helpful for knowledge base question answering) we provide allows not only for better coverage in terms of knowledge base (KB) properties but also for increased output variability in that it permits the generation of multiple questions from the same KB triple.
Kelvin Han, Thiago Castro Ferreira, Claire Gardent
LREC2
2022 MTLens: Machine Translation Output Debugging
abstract
The performance of Machine Translation (MT) systems varies significantly with inputs of diverging features such as topics, genres, and surface properties. Though there are many MT evaluation metrics that generally correlate with human judgments, they are not directly useful in identifying specific shortcomings of MT systems. In this demo, we present a benchmarking interface that enables improved evaluation of specific MT systems in isolation or multiple MT systems collectively by quantitatively evaluating their performance on many tasks across multiple domains and evaluation metrics. Further, it facilitates effective debugging and error analysis of MT output via the use of dynamic filters that help users hone in on problem sentences with specific properties, such as genre, topic, sentence length, etc. The interface can be extended to include additional filters such as lexical, morphological, and syntactic features. Aside from helping debug MT output, it can also help in identifying problems in reference translations and evaluation metrics.
Shreyas Sharma, Kareem Darwish, Lucas Pavanelli, Thiago Castro Ferreira, Mohamed Al-Badrashiny, Kamer Ali Yüksel, Hassan Sawaf
LREC4
2021 Enriching the E2E dataset
abstract
This study introduces an enriched version of the E2E dataset, one of the most popular language resources for data-to-text NLG.We extract intermediate representations for popular pipeline tasks such as discourse ordering, text structuring, lexicalization and referring expression generation, enabling researchers to rapidly develop and evaluate their data-totext pipeline systems.The intermediate representations are extracted by aligning nonlinguistic and text representations through a process called delexicalization, which consists in replacing input referring expressions to entities/attributes with placeholders.The enriched dataset is publicly available.1
Thiago Castro Ferreira, Helena Vaz, Adriana S. Pagano
INLG1
2021 Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation System
abstract
This paper reports results from a reproduction study in which we repeated the human evaluation of the PASS Dutch-language football report generation system (van der Lee et al., 2017).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluations in NLG, in Track A (Paper 1).We aimed to repeat the original study exactly, with the main difference that a different set of evaluators was used.We describe the study design, present the results from the original and the reproduction study, and then compare and analyse the differences between the two sets of results.For the two 'headline' results of average Fluency and Clarity, we find that in both studies, the system was rated more highly for Clarity than for Fluency, and Clarity had higher standard deviation.Clarity and Fluency ratings were higher, and their standard deviations lower, in the reproduction study than in the original study by substantial margins.Clarity had a higher degree of reproducibility than Fluency, as measured by the coefficient of variation.Data and code are publicly available.1
Simon Mille, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001
INLG2
2020 Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing
abstract
This paper introduces the first corpus for Automatic Post-Editing of English and a low-resource language, Brazilian Portuguese. The source English texts were extracted from the WebNLG corpus and automatically translated into Portuguese using a state-of-the-art industrial neural machine translator. Post-edits were then obtained in an experiment with native speakers of Brazilian Portuguese. To assess the quality of the corpus, we performed error analysis and computed complexity indicators measuring how difficult the APE task would be. We report preliminary results of Phrase-Based and Neural Machine Translation Models on this new corpus. Data and code publicly available in our repository.
Felipe Almeida Costa, Thiago Castro Ferreira, Adriana S. Pagano, Wagner Meira Jr.
COLING2
2020 Referring to what you know and do not know: Making Referring Expression Generation Models Generalize To Unseen Entities
abstract
Data-to-text Natural Language Generation (NLG) is the computational process of generating natural language in the form of text or voice from non-linguistic data.A core micro-planning task within NLG is referring expression generation (REG), which aims to automatically generate noun phrases to refer to entities mentioned as discourse unfolds.A limitation of novel REG models is not being able to generate referring expressions to entities not encountered during the training process.To solve this problem, we propose two extensions to NeuralREG, a state-ofthe-art encoder-decoder REG model.The first is a copy mechanism, whereas the second consists of representing the gender and type of the referent as inputs to the model.Drawing on the results of automatic and human evaluation as well as an ablation study using the WebNLG corpus, we contend that our proposal contributes to the generation of more meaningful referring expressions to unseen entities than the original system and related work.Code and all produced data are publicly available 12 .
Rossana Cunha, Thiago Castro Ferreira, Adriana S. Pagano, Fábio Alves
COLING2
2020 DaMata: A Robot-Journalist Covering the Brazilian Amazon Deforestation
abstract
This demo paper introduces DaMata, a robotjournalist covering deforestation in the Brazilian Amazon.The robot-journalist is based on a pipeline architecture of Natural Language Generation, which yields multilingual daily and monthly reports based on the public data provided by DETER, a real-time deforestation satellite monitor developed and maintained by the Brazilian National Institute for Space Research (INPE).DaMata automatically generates reports in Brazilian Portuguese and English and publishes them on the Twitter platform.Corpus and code are publicly available.1
André Luiz Rosa Teixeira, João Campos, Rossana Cunha, Thiago Castro Ferreira, Adriana S. Pagano, Fábio G. Cozman
INLG4
2020 NABU - Multilingual Graph-Based Neural RDF Verbalizer
Diego Moussallem, Dwaraknath Gnaneshwar, Thiago Castro Ferreira, Axel-Cyrille Ngonga Ngomo
ISWC (1)3
2019 Neural data-to-text generation: A comparison between pipeline and end-to-end architectures
abstract
Thiago Castro Ferreira, Chris van der Lee, Emiel van Miltenburg, Emiel Krahmer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Thiago Castro Ferreira, Chris van der Lee, Emiel van Miltenburg, Emiel Krahmer
EMNLP/IJCNLP (1)1
2018 NeuralREG: An end-to-end approach to referring expression generation
abstract
Thiago Castro Ferreira, Diego Moussallem, Ákos Kádár, Sander Wubben, Emiel Krahmer. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Thiago Castro Ferreira, Diego Moussallem, Ákos Kádár, Sander Wubben, Emiel Krahmer
ACL (1)1
2018 Enriching the WebNLG corpus
abstract
This paper describes the enrichment of WebNLG corpus (Gardent et al., 2017a,b), with the aim to further extend its usefulness as a resource for evaluating common NLG tasks, including Discourse Ordering, Lexicalization and Referring Expression Generation.We also produce a silverstandard German translation of the corpus to enable the exploitation of NLG approaches to other languages than English.The enriched corpus is publicly available 1 .
Thiago Castro Ferreira, Diego Moussallem, Emiel Krahmer, Sander Wubben
INLG1
2018 RDF2PT: Generating Brazilian Portuguese Texts from RDF Data
Diego Moussallem, Thiago Castro Ferreira, Marcos Zampieri, Maria Cláudia Cavalcanti, Geraldo Xexéo, Mariana L. Neves, Axel-Cyrille Ngonga Ngomo
LREC2
2017 Generating flexible proper name references in text: Data, models and evaluation
abstract
This study introduces a statistical model able to generate variations of a proper name by taking into account the person to be mentioned, the discourse context and variation.The model relies on the REGnames corpus, a dataset with 53,102 proper name references to 1,000 people in different discourse contexts.We evaluate the versions of our model from the perspective of how human writers produce proper names, and also how human readers process them.The corpus 1 and the model 2 are publicly available.
Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben
EACL (1)1
2017 Linguistic realisation as machine translation: Comparing different MT models for AMR-to-text generation
abstract
In this paper, we study AMR-to-text generation, framing it as a translation task and comparing two different MT approaches (Phrasebased and Neural MT).We systematically study the effects of 3 AMR preprocessing steps (Delexicalisation, Compression, and Linearisation) applied before the MT phase.Our results show that preprocessing indeed helps, although the benefits differ for the two MT models.The implementation of the models are publicly available 1 .
Thiago Castro Ferreira, Iacer Calixto, Sander Wubben, Emiel Krahmer
INLG1
2017 Improving the generation of personalised descriptions
abstract
Referring expression generation (REG) models that use speaker-dependent information require a considerable amount of training data produced by every individual speaker, or may otherwise perform poorly.In this work we propose a simple personalised method for this task, in which speakers are grouped into profiles according to their referential behaviour.Intrinsic evaluation shows that the use of speaker's profiles generally outperforms the personalised method found in previous work.
Thiago Castro Ferreira, Ivandré Paraboni
INLG1
2017 Generating natural language descriptions using speaker-dependent information
abstract
Abstract This paper discusses the issue of human variation in natural language referring expression generation. We introduce a model of content selection that takes speaker-dependent information into account to produce descriptions that closely resemble those produced by each individual, as seen in a number of reference corpora. Results show that our speaker-dependent referring expression generation model outperforms alternatives that do not take human variation into account, or which do so less extensively, and suggest that the use of machine-learning methods may be an ideal approach to mimic complex referential behaviour.
Thiago Castro Ferreira, Ivandré Paraboni
Nat. Lang. Eng.1
2016 Towards more variation in text generation: Developing and evaluating variation models for choice of referential form
abstract
In this study, we introduce a nondeterministic method for referring expression generation. We describe two models that account for individual variation in the choice of referential form in automatically generated text: a Naive Bayes model and a Recurrent Neural Network. Both are evaluated using the VaREG corpus. Then we select the best performing model to generate referential forms in texts from the GREC-2.0 corpus and conduct an evaluation experiment in which humans judge the coherence and comprehensibility of the generated texts, comparing them both with the original references and those produced by a random baseline model.
Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben
ACL (1)1
2016 Task demands and individual variation in referring expressions
abstract
Aiming to improve the human-likeness of natural language generation systems, this study investigates different sources of variation that might influence the production of referring expressions (REs), namely the effect of task demands and inter-intra-individual variation.We collected REs using a discrimination game and varied the instructions, telling speakers that they would get points for being fast, creative, clear, or no incentive would be mentioned.Our results show that taskdemands affected REs production (number of words, number of attributes), and we observe a considerable amount of variation among the length of REs produced by single speakers, as well as among the REs of different speakers referring to the same targets.
Adriana Alexandra Baltaretu, Thiago Castro Ferreira
INLG2
2016 Towards proper name generation: a corpus analysis
abstract
We introduce a corpus for the study of proper name generation.The corpus consists of proper name references to people in webpages, extracted from the Wikilinks corpus.In our analyses, we aim to identify the different ways, in terms of length and form, in which a proper names are produced throughout a text.
Thiago Castro Ferreira, Sander Wubben, Emiel Krahmer
INLG1
2016 Individual Variation in the Choice of Referential Form
abstract
Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Thiago Castro Ferreira, Emiel Krahmer, Sander Wubben
HLT-NAACL1
2014 Classification-Based Referring Expression Generation
Thiago Castro Ferreira, Ivandré Paraboni
CICLing (1)1