VLDB 2026 Research / reviewers in the wild / expert
Aitor Ormazabal
dblp:243/3370
· DBLP profile ↗
7ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 41% Machine translation · 28% Representation and self-supervised learning · 19% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding |
0.9 | 2 | 2021 | Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring · ACL/IJCNLP (1) 2021 Analyzing the Limitations of Cross-lingual Word Embedding Mappings · ACL (1) 2019 |
Natural language and speech › Language models and text generation
multilingual language models |
0.8 | 1 | 2024 | Latxa: An Open Language Model and Evaluation Suite for Basque · ACL (1) 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › source-free domain adaptation
black-box domain adaptation |
0.7 | 1 | 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models · EMNLP 2023 |
Natural language and speech › Machine translation
domain adaptation for machine translation |
0.7 | 1 | 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models · EMNLP 2023 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.7 | 1 | 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models · EMNLP 2023 |
Natural language and speech › Machine translation › monolingual data augmentation
back-translation |
0.6 | 1 | 2022 | Principled Paraphrase Generation with Parallel Corpora · ACL (1) 2022 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.6 | 1 | 2022 | Principled Paraphrase Generation with Parallel Corpora · ACL (1) 2022 |
Natural language and speech › Machine translation
bilingual lexicon induction |
0.3 | 2 | 2021 | Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring · ACL/IJCNLP (1) 2021 Analyzing the Limitations of Cross-lingual Word Embedding Mappings · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning
similarity measure |
0.2 | 1 | 2022 | Principled Paraphrase Generation with Parallel Corpora · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
skip-gram · 0.9language model pretraining · 0.8small white-box language model · 0.7probability-level ensemble · 0.7fine-tuning · 0.7information bottleneck · 0.6adversarial training · 0.6self-learning · 0.5iterative restarts · 0.5linear transformation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal LLMs Do Not Compose Skills Optimally across ModalitiesabstractSkill composition is the ability to combine previously learned skills to solve new tasks. As neural networks acquire increasingly complex skills during their pretraining, it is not clear how successfully they can compose them. In this paper, we focus on Multimodal Large Language Models (MLLM), and study their ability to compose skills across modalities. To this end, we design three evaluation tasks which can be solved sequentially composing two modality-dependent skills, and evaluate several open MLLMs under two main settings: i) prompting the model to directly solve the task, and ii) using a two-step cascaded inference approach, which manually enforces the composition of the two skills for a given task. Even with these straightforward compositions, we find that all evaluated MLLMs exhibit a significant cross-modality skill composition gap. To mitigate the aforementioned gap, we explore two alternatives: i) use chain-of-thought prompting to explicitly instruct MLLMs for skill composition and ii) a specific fine-tuning recipe to promote skill composition. Although those strategies improve model performance, they still exhibit significant skill composition gaps, suggesting that more research is needed to improve cross-modal skill composition in MLLMs. Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune |
LREC | 2 |
| 2025 | Improving the Efficiency of Visually Augmented Language ModelsabstractDespite the impressive performance of autoregressive Language Models (LM) it has been shown that due to reporting bias, LMs lack visual knowledge, i.e. they do not know much about the visual world and its properties. To augment LMs with visual knowledge, existing solutions often rely on explicit images, requiring time-consuming retrieval or image generation systems. This paper shows that explicit images are not necessary to visually augment an LM. Instead, we use visually-grounded text representations obtained from the well-known CLIP multimodal system. For a fair comparison, we modify VALM, a visually-augmented LM which uses image retrieval and representation, to work directly with visually-grounded text representations. We name this new model BLIND-VALM. We show that BLIND-VALM performs on par with VALM for Visual Language Understanding (VLU), Natural Language Understanding (NLU) and Language Modeling tasks, despite being significantly more efficient and simpler. We also show that scaling up our model within the compute budget of VALM, either increasing the model or pre-training corpus size, we outperform VALM for all the evaluation tasks. Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune |
COLING | 2 |
| 2024 | Latxa: An Open Language Model and Evaluation Suite for BasqueabstractJulen Etxaniz, Oscar Sainz, Naiara Perez, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Julen Etxaniz, Oscar Sainz, Naiara Miguel, Itziar Aldabe, German Rigau, Eneko Agirre, Aitor Ormazabal, Mikel Artetxe, Aitor Soroa |
ACL (1) | 7 |
| 2023 | CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsabstractMethods for adapting language models (LMs) to new tasks and domains have traditionally assumed white-box access to the model, and work by modifying its parameters.However, this is incompatible with a recent trend in the field, where the highest quality models are only available as black-boxes through inference APIs.Even when the model weights are available, the computational cost of fine-tuning large LMs can be prohibitive for most practitioners.In this work, we present a lightweight method for adapting large LMs to new domains and tasks, assuming no access to their weights or intermediate activations.Our approach fine-tunes a small white-box LM and combines it with the large black-box LM at the probability level through a small network, learned on a small validation set.We validate our approach by adapting a large LM (OPT-30B) to several domains and a downstream task (machine translation), observing improved performance in all cases, of up to 9%, while using a domain expert 23x smaller. Aitor Ormazabal, Mikel Artetxe, Eneko Agirre |
EMNLP | 1 |
| 2022 | Principled Paraphrase Generation with Parallel CorporaabstractRound-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision.In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to nonparaphrase pairs sharing a single ambiguous translation.Based on these insights, we design an alternative similarity metric that mitigates this issue by requiring the entire translation distribution to match, and implement a relaxation of it through the Information Bottleneck method.Our approach incorporates an adversarial term into MT training in order to learn representations that encode as much information about the reference translation as possible, while keeping as little information about the input as possible.Paraphrases can be generated by decoding back to the source from this representation, without having to generate pivot translations.In addition to being more principled and efficient than round-trip MT, our approach offers an adjustable parameter to control the fidelity-diversity trade-off, and obtains better results in our experiments. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL (1) | 1 |
| 2021 | Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context AnchoringabstractRecent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes departures from this assumption. In this paper, we propose an alternative approach that does not have this limitation, while requiring a weak seed dictionary (e.g., a list of identical words) as the only form of supervision. Rather than aligning two fixed embedding spaces, our method works by fixing the target language embeddings, and learning a new set of embeddings for the source language that are aligned with them. To that end, we use an extension of skip-gram that leverages translated context words as anchor points, and incorporates self-learning and iterative restarts to reduce the dependency on the initial dictionary. Our approach outperforms conventional mapping methods on bilingual lexicon induction, and obtains competitive results in the downstream XNLI task. Aitor Ormazabal, Mikel Artetxe, Aitor Soroa, Gorka Labaka, Eneko Agirre |
ACL/IJCNLP (1) | 1 |
| 2019 | Analyzing the Limitations of Cross-lingual Word Embedding MappingsabstractRecent research in cross-lingual word embeddings has almost exclusively focused on offline methods, which independently train word embeddings in different languages and map them to a shared space through linear transformations.While several authors have questioned the underlying isomorphism assumption, which states that word embeddings in different languages have approximately the same structure, it is not clear whether this is an inherent limitation of mapping approaches or a more general issue when learning crosslingual embeddings.So as to answer this question, we experiment with parallel corpora, which allows us to compare offline mapping to an extension of skip-gram that jointly learns both embedding spaces.We observe that, under these ideal conditions, joint learning yields to more isomorphic embeddings, is less sensitive to hubness, and obtains stronger results in bilingual lexicon induction.We thus conclude that current mapping methods do have strong limitations, calling for further research to jointly learn cross-lingual embeddings with a weaker cross-lingual signal. Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, Eneko Agirre |
ACL (1) | 1 |