VLDB 2026 Research / reviewers in the wild / expert
Jeremy Barnes 0001
dblp:70/6331-1
· DBLP profile ↗
17ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-8043-8058ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditioning LLMs to Generate Code-Switched Text
Maite Heredia, Gorka Labaka, Jeremy Barnes 0001, Aitor Soroa |
LREC | 3 |
| 2026 | Benchmarking Mathematical Reasoning in a Low-Resource Language: Structured Prompting and Evaluation in Basque
Inigo Martinez-Criado, Aitor Soroa, Jeremy Barnes 0001 |
LREC | 3 |
| 2025 | Truth Knows No Language: Evaluating Truthfulness Beyond EnglishabstractBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes, Pablo Gamallo, Iria de-Dios-Flores, Rodrigo Agerri. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Blanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes 0001, Pablo Gamallo 0001, Iria de-Dios-Flores, Rodrigo Agerri |
ACL (1) | 4 |
| 2025 | IberoBench: A Benchmark for LLM Evaluation in Iberian LanguagesabstractThe current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present IberoBench, a multilingual, multi-task benchmark for Iberian languages (i.e., Basque, Catalan, Galician, European Spanish and European Portuguese) built on the LM Evaluation Harness framework. The benchmark consists of 62 tasks divided into 179 subtasks. We evaluate 33 existing LLMs on IberoBench on 0- and 5-shot settings. We also explore the issues we encounter when working with the Harness and our approach to solving them to ensure high-quality evaluation. Irene Baucells de la Peña, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Pérez, Anna Salles, Susana Sotelo Docío, Júlia Falcão, José Javier Saiz, Robiert Sepúlveda-Torres, Jeremy Barnes 0001, Pablo Gamallo 0001, Aitor Gonzalez-Agirre, German Rigau, Marta Villegas |
COLING | 11 |
| 2024 | XNLIeu: a dataset for cross-lingual NLI in BasqueabstractMaite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes, Aitor Soroa. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Maite Heredia, Julen Etxaniz, Muitze Zulaika, Xabier Saralegi, Jeremy Barnes 0001, Aitor Soroa |
NAACL-HLT | 5 |
| 2021 | Structured Sentiment Analysis as Dependency Graph ParsingabstractJeremy Barnes, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jeremy Barnes 0001, Robin Kurtz, Stephan Oepen, Lilja Øvrelid, Erik Velldal |
ACL/IJCNLP (1) | 1 |
| 2021 | Evaluating morphological typology in zero-shot cross-lingual transferabstractAntonio Martínez-García, Toni Badia, Jeremy Barnes. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Antonio Martínez-García, Toni Badia, Jeremy Barnes 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | If you've got it, flaunt it: Making the most of fine-grained sentiment annotationsabstractFine-grained sentiment analysis attempts to extract sentiment holders, targets and polar expressions and resolve the relationship between them, but progress has been hampered by the difficulty of annotation.Targeted sentiment analysis, on the other hand, is a more narrow task, focusing on extracting sentiment targets and classifying their polarity.In this paper, we explore whether incorporating holder and expression information can improve target extraction and classification and perform experiments on eight English datasets.We conclude that jointly predicting target and polarity BIO labels improves target extraction, and that augmenting the input text with gold expressions generally improves targeted polarity classification.This highlights the potential importance of annotating expressions for fine-grained sentiment datasets.At the same time, our results show that performance of current models for predicting polar expressions is poor, hampering the benefit of this information in practice. Jeremy Barnes 0001, Lilja Øvrelid, Erik Velldal |
EACL | 1 |
| 2021 | Multi-task Learning of Negation and Speculation for Targeted Sentiment ClassificationabstractThe majority of work in targeted sentiment analysis has concentrated on finding better methods to improve the overall results.Within this paper we show that these models are not robust to linguistic phenomena, specifically negation and speculation.In this paper, we propose a multi-task learning method to incorporate information from syntactic and semantic auxiliary tasks, including negation and speculation scope detection, to create English-language models that are more robust to these phenomena.Further we create two challenge datasets to evaluate model performance on negated and speculative samples.We find that multi-task models and transfer learning via language modelling can improve performance on these challenge datasets, but the overall performances indicate that there is still much room for improvement.We release both the datasets and the source code at https://github.com/ jerbarnes/multitask_negation_ for_targeted_sentiment. Jeremy Barnes 0001 |
NAACL-HLT | 2 |
| 2021 | Improving sentiment analysis with multi-task learning of negationabstractAbstract Sentiment analysis is directly affected by compositional phenomena in language that act on the prior polarity of the words and phrases found in the text.Negationis the most prevalent of these phenomena, and in order to correctly predict sentiment, a classifier must be able to identify negation and disentangle the effect that its scope has on the final polarity of a text. This paper proposes a multi-task approach to explicitly incorporate information about negation in sentiment analysis, which we show outperforms learning negation implicitly in an end-to-end manner. We describe our approach, a cascading and hierarchical neural architecture with selective sharing of Long Short-term Memory layers, and show that explicitly training the model with negation as an auxiliary task helps improve the main task of sentiment analysis. The effect is demonstrated across several different standard English-language data sets for both tasks, and we analyze several aspects of our system related to its performance, varying types and amounts of input data and different multi-task setups. Jeremy Barnes 0001, Erik Velldal, Lilja Øvrelid |
Nat. Lang. Eng. | 1 |
| 2020 | Named Entity Recognition without Labelled Data: A Weak Supervision ApproachabstractNamed Entity Recognition (NER) performance often degrades rapidly when applied to target domains that differ from the texts observed during training.When in-domain labelled data is available, transfer learning techniques can be used to adapt existing NER models to the target domain.But what should one do when there is no hand-labelled data for the target domain?This paper presents a simple but powerful approach to learn NER models in the absence of labelled data through weak supervision.The approach relies on a broad spectrum of labelling functions to automatically annotate texts from the target domain.These annotations are then merged together using a hidden Markov model which captures the varying accuracies and confusions of the labelling functions.A sequence labelling model can finally be trained on the basis of this unified annotation.We evaluate the approach on two English datasets (CoNLL 2003 and news articles from Reuters and Bloomberg) and demonstrate an improvement of about 7 percentage points in entity-level F 1 scores compared to an out-of-domain neural NER model. Pierre Lison, Jeremy Barnes 0001, Aliaksandr Hubin, Samia Touileb |
ACL | 2 |
| 2020 | A Fine-grained Sentiment Dataset for NorwegianabstractWe here introduce NoReC_fine, a dataset for fine-grained sentiment analysis in Norwegian, annotated with respect to polar expressions, targets and holders of opinion. The underlying texts are taken from a corpus of professionally authored reviews from multiple news-sources and across a wide variety of domains, including literature, games, music, products, movies and more. We here present a detailed description of this annotation effort. We provide an overview of the developed annotation guidelines, illustrated with examples and present an analysis of inter-annotator agreement. We also report the first experimental results on the dataset, intended as a preliminary benchmark for further experiments. Lilja Øvrelid, Petter Mæhlum, Jeremy Barnes 0001, Erik Velldal |
LREC | 3 |
| 2019 | Embedding Projection for Targeted Cross-lingual Sentiment: Model Comparisons and a Real-World StudyabstractSentiment analysis benefits from large, hand-annotated resources in order to train and test machine learning models, which are often data hungry. While some languages, e.g., English, have a vast arrayof these resources, most under-resourced languages do not, especially for fine-grained sentiment tasks, such as aspect-level or targeted sentiment analysis. To improve this situation, we propose a cross-lingual approach to sentiment analysis that is applicable to under-resourced languages and takes into account target-level information. This model incorporates sentiment information into bilingual distributional representations, byjointly optimizing them for semantics and sentiment, showing state-of-the-art performance at sentence-level when combined with machine translation. The adaptation to targeted sentiment analysis on multiple domains shows that our model outperforms other projection-based bilingual embedding methods on binary targetedsentiment tasks. Our analysis on ten languages demonstrates that the amount of unlabeled monolingual data has surprisingly little effect on the sentiment results. As expected, the choice of a annotated source language for projection to a target leads to better results for source-target language pairs which are similar. Therefore, our results suggest that more efforts should be spent on the creation of resources for less similar languages tothose which are resource-rich already. Finally, a domain mismatch leads to a decreased performance. This suggests resources in any language should ideally cover varieties of domains. Jeremy Barnes 0001, Roman Klinger |
J. Artif. Intell. Res. | 1 |
| 2018 | Bilingual Sentiment Embeddings: Joint Projection of Sentiment Across LanguagesabstractSentiment analysis in low-resource languages suffers from a lack of annotated corpora to estimate high-performing models.Machine translation and bilingual word embeddings provide some relief through cross-lingual sentiment approaches.However, they either require large amounts of parallel data or do not sufficiently capture sentiment information.We introduce Bilingual Sentiment Embeddings (BLSE), which jointly represent sentiment information in a source and target language.This model only requires a small bilingual lexicon, a source-language corpus annotated for sentiment, and monolingual word embeddings for each language.We perform experiments on three language combinations (Spanish, Catalan, Basque) for sentencelevel cross-lingual sentiment classification and find that our model significantly outperforms state-of-the-art methods on four out of six experimental setups, as well as capturing complementary information to machine translation.Our analysis of the resulting embedding space provides evidence that it represents sentiment information in the resource-poor target language without any annotated data in that language. Jeremy Barnes 0001, Roman Klinger, Sabine Schulte im Walde |
ACL (1) | 1 |
| 2018 | Projecting Embeddings for Domain Adaption: Joint Modeling of Sentiment Analysis in Diverse DomainsabstractDomain adaptation for sentiment analysis is challenging due to the fact that supervised classifiers are very sensitive to changes in domain. The two most prominent approaches to this problem are structural correspondence learning and autoencoders. However, they either require long training times or suffer greatly on highly divergent domains. Inspired by recent advances in cross-lingual sentiment analysis, we provide a novel perspective and cast the domain adaptation problem as an embedding projection task. Our model takes as input two mono-domain embedding spaces and learns to project them to a bi-domain space, which is jointly optimized to (1) project across domains and to (2) predict sentiment. We perform domain adaptation experiments on 20 source-target domain pairs for sentiment classification and report novel state-of-the-art results on 11 domain pairs, including the Amazon domain adaptation datasets and SemEval 2013 and 2016 datasets. Our analysis shows that our model performs comparably to state-of-the-art approaches on domains that are similar, while performing significantly better on highly divergent domains. Our code is available at https://github.com/jbarnesspain/domain_blse Jeremy Barnes 0001, Roman Klinger, Sabine Schulte im Walde |
COLING | 1 |
| 2018 | MultiBooked: A Corpus of Basque and Catalan Hotel Reviews Annotated for Aspect-level Sentiment Classification
Jeremy Barnes 0001, Toni Badia, Patrik Lambert |
LREC | 1 |
| 2016 | Exploring Distributional Representations and Machine Translation for Aspect-based Cross-lingual Sentiment ClassificationabstractCross-lingual sentiment classification (CLSC) seeks to use resources from a source language in order to detect sentiment and classify text in a target language. Almost all research into CLSC has been carried out at sentence and document level, although this level of granularity is often less useful. This paper explores methods for performing aspect-based cross-lingual sentiment classification (aspect-based CLSC) for under-resourced languages. Given the limited nature of parallel data for many languages, we would like to make the most of this resource for our task. We compare zero-shot learning, bilingual word embeddings, stacked denoising autoencoder representations and machine translation techniques for aspect-based CLSC. Each of these approaches requires differing amounts of parallel data. We show that models based on distributed semantics can achieve comparable results to machine translation on aspect-based CLSC and give an analysis of the errors found for each method. Jeremy Barnes 0001, Patrik Lambert, Toni Badia |
COLING | 1 |