António Branco

dblp:51/1654 · also António Horta Branco · DBLP profile ↗
← Back
51ranked-venue papers
17as first author
9since 2021 · last 2026
0000-0002-7174-4942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 17 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021
YearPublicationVenuePosition
2026 ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
Ricardo Campos 0001, Raquel Sequeira, Sara Nerea, Inês Cantante, Diogo Folques, Luís Filipe Cunha, João Canavilhas, António Branco, Alípio Mário Jorge, Sérgio Nunes 0001, Nuno Guimarães, Purificação Silvano
ECIR (4)8
2026 pt-image-ir-dataset: An Image Retrieval Dataset in European Portuguese
Rodrigo Duarte, António Branco, Hugo Proença 0001, Ricardo Campos 0001
ECIR (4)2
2026 ImageSeek: A Hybrid Text-to-Image Image Retrieval System for Domain-Specific Collections
Rodrigo Duarte, António Branco, Hugo Proença 0001, Ricardo Campos 0001
ECIR (4)3
2026 Sovereign AI-based Public Services Are Viable and Affordable
António Branco, Luis M. S. Gomes, Rodrigo Santos, Eduardo Santos, João Silva 0004, Madalena Rodrigues
LREC1
2025 Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
abstract
Instruction-guided image editing consists in taking an image and an instruction and delivering that image altered according to that instruction. State-of-the-art approaches to this task suffer from the typical scaling up and domain adaptation hindrances related to supervision as they eventually resort to some kind of task-specific labelling, masking or training. We propose a novel approach that does without any such task-specific supervision and offers thus a better potential for improvement. Its assessment demonstrates that it is highly effective, achieving very competitive performance.
Rodrigo Santos, António Branco, João Silva 0004, João Rodrigues 0001
COLING2
2022 Transferring Confluent Knowledge to Argument Mining
abstract
Relevant to all application domains where it is important to get at the reasons underlying sentiments and decisions, argument mining seeks to obtain structured arguments from unstructured text and has been addressed by approaches typically involving some feature and/or neural architecture engineering. By adopting a transfer learning methodology, and by means of a systematic study with a wide range of knowledge sources promisingly suitable to leverage argument mining, the aim of this paper is to empirically assess the potential of transferring such knowledge learned with confluent tasks. By adopting a lean approach that dispenses with heavier feature and model engineering, this study permitted both to gain novel empirically based insights into the argument mining task and to establish new state of the art levels of performance for its three main sub-tasks, viz. identification of argument components, classification of the components, and determination of the relation among them.
João Rodrigues 0001, António Branco
COLING2
2022 Language Driven Image Editing via Transformers
abstract
With the emergence of specifically tailored neural architectures that cope with both modalities, cross-modal language and image processing has attracted increasing attention. A major motivation has been the search for a quantum leap in language understanding supported by visual grounding, which has been oriented mostly to solve tasks where language descriptions of images are to be provided, and vice-versa, where images are to be generated on the basis of keywords. Adopting a distinct angle of inquiry, this paper addresses rather the cross-modal challenge of language driven image design, focusing on the task of editing an image on the basis of language instructions to modify it. And adopting as well a distinct research path, which dispenses with specifically tailored architectures, the approach proposed here resorts rather to a general purpose, suitably instantiated neural architecture of the Transformer class. Experimentation with this approach delivered very encouraging results, empirically demonstrating that this is an effective methodology for language driven image design and the basis for further advances in cross-modal processing and its applications with affordable compute and data.
Rodrigo Santos, António Branco, João Silva 0004
ICTAI2
2022 Universal Grammatical Dependencies for Portuguese with CINTIL Data, LX Processing and CLARIN support
abstract
The grammatical framework for the mapping between linguistic form and meaning representation known as Universal Dependencies relies on a non-constituency syntactic analysis that is centered on the notion of grammatical relation (e.g. Subject, Object, etc.). Given its core goal of providing a common set of analysis primitives suitable to every natural language, and its practical objective of fostering their computational grammatical processing, it keeps being an active domain of research in science and technology of language. This paper presents a new collection of quality language resources for the computational processing of the Portuguese language under the Universal Dependencies framework (UD). This is an all-encompassing, publicly available open collection of mutually consistent and inter-operable scientific resources that includes reliably annotated corpora, top-performing processing tools and expert support services: a new UPOS-annotated corpus, CINTIL-UPos, with 675K tokens and a new UD treebank, CINTIL-UDep Treebank, with nearly 38K sentences; a UPOS tagger, LX-UTagger, and a UD parser, LX-UDParser, trained on these corpora, available both as local stand-alone tools and as remote web-based services; and helpdesk support ensured by the Knowledge Center for the Science and Technology of Portuguese of the CLARIN research infrastructure.
António Branco, João Silva 0004, Luís Gomes 0002, João Rodrigues 0001
LREC1
2021 Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning
abstract
Commonsense is a quintessential human capacity that has been a core challenge to Artificial Intelligence since its inception.Impressive results in Natural Language Processing tasks, including in commonsense reasoning, have consistently been achieved with Transformer neural language models, even matching or surpassing human performance in some benchmarks.Recently, some of these advances have been called into question: so called data artifacts in the training data have been made evident as spurious correlations and shallow shortcuts that in some cases are leveraging these outstanding results.In this paper we seek to further pursue this analysis into the realm of commonsense related language processing tasks.We undertake a study on different prominent benchmarks that involve commonsense reasoning, along a number of key stress experiments, thus seeking to gain insight on whether the models are learning transferable generalizations intrinsic to the problem at stake or just taking advantage of incidental shortcuts in the data items.The results obtained indicate that most datasets experimented with are problematic, with models resorting to non-robust features and appearing not to be learning and generalizing towards the overall tasks intended to be conveyed or exemplified by the datasets.Question: What is something someone driving a car needs even to begin?Answer A: practice.Answer B: feet.
Ruben Branco, António Branco, João Rodrigues 0001, João Silva 0004
EMNLP (1)2
2020 Comparative Probing of Lexical Semantics Theories for Cognitive Plausibility and Technological Usefulness
abstract
Lexical semantics theories differ in advocating that the meaning of words is represented as an inference graph, a feature mapping or a vector space, thus raising the question: is it the case that one of these approaches is superior to the others in representing lexical semantics appropriately?Or in its non antagonistic counterpart: could there be a unified account of lexical semantics where these approaches seamlessly emerge as (partial) renderings of (different) aspects of a core semantic knowledge base?In this paper, we contribute to these research questions with a number of experiments that systematically probe different lexical semantics theories for their levels of cognitive plausibility and of technological usefulness.The empirical findings obtained from these experiments advance our insight on lexical semantics as the feature-based approach emerges as superior to the other ones, and arguably also move us closer to finding answers to the research questions above.
António Branco, João Rodrigues 0001, Malgorzata Salawa, Ruben Branco, Chakaveh Saedi
COLING1
2020 A Shared Task of a New, Collaborative Type to Foster Reproducibility: A First Exercise in the Area of Language Science and Technology with REPROLANG2020
abstract
n this paper, we introduce a new type of shared task — which is collaborative rather than competitive — designed to support and fosterthe reproduction of research results. We also describe the first event running such a novel challenge, present the results obtained, discussthe lessons learned and ponder on future undertakings.
António Branco, Nicoletta Calzolari, Piek Vossen, Gertjan van Noord, Dieter Van Uytvanck, João Silva 0004, Luís Gomes 0002, Willem Elbers
LREC1
2020 The MWN.PT WordNet for Portuguese: Projection, Validation, Cross-lingual Alignment and Distribution
abstract
The objective of the present paper is twofold, to present the MWN.PT WordNet and to report on its construction and on the lessons learned with it. The MWN.PT WordNet for Portuguese includes 41,000 concepts, expressed by 38,000 lexical units. Its synsets were manually validated and are linked to semantically equivalent synsets of the Princeton WordNet of English, and thus transitively to the many wordnets for other languages that are also linked to this English wordnet. To the best of our knowledge, it is the largest high quality, manually validated and cross-lingually integrated, wordnet of Portuguese distributed for reuse. Its construction was initiated more than one decade ago and its description is published for the first time in the present paper. It follows a three step <projection, validation with alignment, completion> methodology consisting on the manual validation and expansion of the outcome of an automatic projection procedure of synsets and their hypernym relations, followed by another automatic procedure that transferred the relations of remaining semantic types across wordnets of different languages.
António Branco, Sara Grilo, Márcia Bolrinha, Chakaveh Saedi, Ruben Branco, João Silva 0004, Andreia Querido, Rita de Carvalho, Rosa Del Gaudio, Mariana Avelãs, Clara Pinto
LREC1
2020 The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology
abstract
This paper presents the BDCamões Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese that in its inaugural version includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres, covering a time span from the 16th to the 21st century, and adhering to different orthographic conventions. Many of the texts in the corpus have also been automatically parsed with state-of-the-art language processing tools, forming the BDCamões Treebank subcorpus. This set of characteristics makes of BDCamões an invaluable resource for research in language technology (e.g. authorship detection, genre classification, etc.) and in language science and digital humanities (e.g. comparative literature, diachronic linguistics, etc.).
Sara Grilo, Márcia Bolrinha, João Silva 0004, Rui Vaz, António Branco
LREC5
2020 The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
abstract
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
LREC24
2020 Reproduction and Revival of the Argument Reasoning Comprehension Task
abstract
Reproduction of scientific findings is essential for scientific development across all scientific disciplines and reproducing results of previous works is a basic requirement for validating the hypothesis and conclusions put forward by them. This paper reports on the scientific reproduction of several systems addressing the Argument Reasoning Comprehension Task of SemEval2018. Given a recent publication that pointed out spurious statistical cues in the data set used in the shared task, and that produced a revised version of it, we also evaluated the reproduced systems with this new data set. The exercise reported here shows that, in general, the reproduction of these systems is successful with scores in line with those reported in SemEval2018. However, the performance scores are worst than those, and even below the random baseline, when the reproduced systems are run over the revised data set expunged from data artifacts. This demonstrates that this task is actually a much harder challenge than what could have been perceived from the inflated, close to human-level performance scores obtained with the data set used in SemEval2018. This calls for a revival of this task as there is much room for improvement until systems may come close to the upper bound provided by human performance.
João Rodrigues 0001, Ruben Branco, João Silva 0004, António Branco
LREC4
2019 Assessing Wordnets with WordNet Embeddings
abstract
An effective conversion method was proposed in the literature to obtain a lexical semantic space from a lexical semantic graph, thus permitting to obtain Word-Net embeddings from WordNets.In this paper, we propose the exploitation of this conversion methodology as the basis for the comparative assessment of WordNets: given two WordNets, their relative quality in terms of capturing the lexical semantics of a given language, can be assessed by (i) converting each WordNet into the corresponding semantic space (i.e.into WordNet embeddings), (ii) evaluating the resulting WordNet embeddings under the typical semantic similarity prediction task used to evaluate word embeddings in general; and (iii) comparing the performance in that task of the two word embeddings, extracted from the two WordNets.A better performance in that evaluation task results from the word embeddings that are better at capturing the semantic similarity of words, which, in turn, result from the WordNet that is of higher quality at capturing the semantics of words.
Ruben Branco, João Rodrigues 0001, Chakaveh Saedi, António Branco
GWC4
2018 Attention Focusing for Neural Machine Translation by Bridging Source and Target Embeddings
abstract
In neural machine translation, a source sequence of words is encoded into a vector from which a target sequence is generated in the decoding phase.Differently from statistical machine translation, the associations between source words and their possible target counterparts are not explicitly stored.Source and target words are at the two ends of a long information processing procedure, mediated by hidden states at both the source encoding and the target decoding phases.This makes it possible that a source word is incorrectly translated into a target word that is not any of its admissible equivalent counterparts in the target language.
Shaohui Kuang, António Branco, Weihua Luo, Deyi Xiong
ACL (1)3
2018 ELRI - European Language Resources Infrastructure
abstract
We describe the European Language Resources Infrastructure project, whose main aim is the provision of an infrastructure to help collect, prepare and share language resources that can in turn improve translation services in Europe.
Thierry Etchegoyhen, Borja Anza Porras, Andoni Azpeitia, Eva Martínez Garcia, Paulo Vale, José Luis Fonseca, Teresa Lynn, Jane Dunne, Federico Gaspari, Andy Way, Victoria Arranz, Khalid Choukri, Vladimir Popescu, Pedro Neiva, Rui Neto, Maite Melero, David Pérez-Fernández, António Branco, Ruben Branco, Luís Gomes 0002
EAMT18
2018 Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese
João Rodrigues 0001, António Branco
LREC2
2018 Semantic Equivalence Detection: Are Interrogatives Harder than Declaratives?
João Rodrigues 0001, Chakaveh Saedi, António Branco, João Silva 0004
LREC3
2018 We Are Depleting Our Research Subject as We Are Investigating It: In Language Technology, more Replication and Diversity Are Needed
António Branco
LREC1
2018 Browsing and Supporting Pluricentric Global Wordnet, or just your Wordnet of Interest
António Branco, Ruben Branco, Chakaveh Saedi, João Silva 0004
LREC1
2018 WordnetLoom - a Multilingual Wordnet Editing System Focused on Graph-based Presentation
abstract
The paper presents a new re-built and expanded, version 2.0 of WordnetLoom -an open wordnet editor.It facilitates work on a multilingual system of wordnets, is based on efficient software architecture of thin client, and offers more flexibility in enriching wordnet representation.This new version is built on the experience collected during the use of the previous one for more than 10 years of plWordNet development.We discuss its extensions motivated by the collected experience.A special focus is given to the development of a variant for the needs of MultiWordnet of Portuguese, which is based on a very different wordnet development model.
Tomasz Naskret, Agnieszka Dziob, Maciej Piasecki, Chakaveh Saedi, António Branco
GWC5
2016 CINTIL DependencyBank PREMIUM - A Corpus of Grammatical Dependencies for Portuguese
Rita de Carvalho, Andreia Querido, Marisa Campos, Rita Valadas Pereira, João Silva 0004, António Branco
LREC6
2016 Evaluating Machine Translation in a Usage Scenario
Rosa Del Gaudio, Aljoscha Burchardt, António Branco
LREC3
2016 Word Sense-Aware Machine Translation: Including Senses as Contextual Features for Improved Translation Models
Steven Neale, Luís Gomes 0002, Eneko Agirre, Oier Lopez de Lacalle, António Branco
LREC5
2016 QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages
Arantxa Otegi, Nora Aranberri, António Branco, Jan Hajic 0001, Martin Popel, Kiril Ivanov Simov, Eneko Agirre, Petya Osenova, Rita Valadas Pereira, João Silva 0004, Steven Neale
LREC3
2016 Bootstrapping a Hybrid MT System to a New Language Pair
João Rodrigues 0001, Nuno Rendeiro, Andreia Querido, Sanja Stajner, António Branco
LREC5
2016 Use of Domain-Specific Language Resources in Machine Translation
Sanja Stajner, Andreia Querido, Nuno Rendeiro, João Rodrigues 0001, António Branco
LREC5
2014 Answering List Questions using Web as a corpus
abstract
This paper supports the demo of LX-ListQuestion, a Web Question Answering System that exploit the redundancy of in-formation available in the Web to answer List Questions in the form of Word Cloud. 1
Patrícia Nunes Gonçalves, António Branco
EACL2
2014 The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite
LREC7
2014 Coping with highly imbalanced datasets: A case study with definition extraction in a multilingual setting
abstract
Abstract This paper addresses the task of automatic extraction of definitions by thoroughly exploring an approach that solely relies on machine learning techniques, and by focusing on the issue of the imbalance of relevant datasets. We obtained a breakthrough in terms of the automatic extraction of definitions, by extensively and systematically experimenting with different sampling techniques and their combination, as well as a range of different types of classifiers. Performance consistently scored in the range of 0.95–0.99 of area under the receiver operating characteristics, with a notorious improvement between 17 and 22 percentage points regarding the baseline of 0.73–0.77, for datasets with different rates of imbalance. Thus, the present paper also represents a contribution to the seminal work in natural language processing that points toward the importance of exploring the research path of applying sampling techniques to mitigate the bias induced by highly imbalanced datasets, and thus greatly improving the performance of a large range of tools that rely on them.
Rosa Del Gaudio, Gustavo Batista, António Branco
Nat. Lang. Eng.3
2013 Compressing Multi-document Summaries through Sentence Simplification
Sara Silveira, António Branco
ICAART (2)2
2012 Aspectual Type and Temporal Relation Classification
Francisco Costa, António Branco
EACL2
2012 A PropBank for Portuguese: the CINTIL-PropBank
António Branco, Catarina Carvalheiro, Sílvia Pereira, Sara Silveira, João Silva 0004, Sérgio Castro, João Graça
LREC1
2012 TimeBankPT: A TimeML Annotated Corpus of Portuguese
Francisco Costa, António Branco
LREC2
2012 Treebanking by Sentence and Tree Transformation: Building a Treebank to support Question Answering in Portuguese
Patrícia Nunes Gonçalves, Rita Santos, António Branco
LREC3
2012 Extracting Multi-document Summaries with a Double Clustering Approach
Sara Silveira, António Branco
NLDB2
2010 Temporal Information Processing of a New Language: Fast Porting with Minimal Resources
Francisco Costa, António Branco
ACL2
2010 Developing a Deep Linguistic Databank Supporting a Collection of Treebanks: the CINTIL DeepGramBank
António Branco, Francisco Costa, João Silva 0004, Sara Silveira, Sérgio Castro, Mariana Avelãs, Clara Pinto, João Graça
LREC1
2010 Top-Performing Robust Constituency Parsing of Portuguese: Freely Available in as Many Ways as you Can Get it
João Silva 0004, António Branco, Patrícia Nunes Gonçalves
LREC2
2008 LX-Service: Web Services of Language Technology for Portuguese
António Branco, Francisco Costa, Pedro Martins 0002, Filipe Nunes, João Silva 0004, Sara Silveira
LREC1
2008 Anaphora Resolution Exercise: an Overview
Constantin Orasan, Dan Cristea, Ruslan Mitkov, António Branco
LREC4
2006 A Suite of Shallow Processing Tools for Portuguese: LX-Suite
António Branco, João Silva 0004
EACL1
2006 Open Resources and Tools for the Shallow Processing of Portuguese: The TagShare Project
Florbela Barreto, António Branco, Amália Mendes, Maria Fernanda Bacelar do Nascimento, Filipe Nunes, João Silva 0004
LREC2
2004 Evaluating Solutions for the Rapid Development of State-of-the-Art POS Taggers for Portuguese
António Branco, João Silva 0004
LREC1
2002 Nexing Corpus: a corpus of verbal protocols on syllogistic reasoning
António Branco, José Leitão, João Silva 0004, Luís Gomes 0003
LREC1
2002 Binding Machines
abstract
Binding constraints form one of the most robust modules of grammatical knowledge. Despite their crosslinguistic generality and practical relevance for anaphor resolution, they have resisted full integration into grammar processing. The ultimate reason for this is to be found in the original exhaustive coindexation rationale for their specification and verification. As an alternative, we propose an approach which, while permitting a unification-based specification of binding constraints, allows for a verification methodology that helps to overcome previous drawbacks. This alternative approach is based on the rationale that anaphoric nominals can be viewed as binding machines.
António Branco
Comput. Linguistics1
2000 Binding Constraints as Instructions of Binding Machines
António Branco
COLING1
1996 Branching Split Obliqueness at the Syntax-Semantics Interface
António Branco
COLING1
1996 Subject-oriented and non Subject-oriented Long-distance Anaphora : an Integrated Approach
António Branco, Palmira Marrafa
PACLIC1