Gerasimos Lampouras

dblp:99/1688 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 38% Efficient and distributed learning · 29% Transfer learning and domain adaptation · 13%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 67% Image and video processing · 33%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference acceleration
draft model
0.912025
Mixture of Attentions For Speculative Decoding · ICLR 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Mixture of Attentions For Speculative Decoding · ICLR 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.912025
Human-inspired Episodic Memory for Infinite Context LLMs · ICLR 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Human-inspired Episodic Memory for Infinite Context LLMs · ICLR 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
Mixture of Attentions For Speculative Decoding · ICLR 2025
Computer vision › Segmentation and scene understanding
scene understanding
0.812024
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024
Visual content generation and editing
image editing
0.812024
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024
Image and video processing › image decomposition › image separation
layer separation
0.812024
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024
Visual content generation and editing › image generation
text-to-image generation
0.812024
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.612022
Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022
Machine learning › Learning paradigms
curriculum learning
0.612022
Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
zero-shot cross-lingual transfer
0.612022
Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022
Natural language and speech › Language models and text generation › text generation › data-to-text generation
concept-to-text generation
0.512021
Generalising Multilingual Concept-to-Text NLG with Language Agnostic Delexicalisation · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text generation
multilingual text generation
0.512021
Generalising Multilingual Concept-to-Text NLG with Language Agnostic Delexicalisation · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
text generation
0.512021
Generalising Multilingual Concept-to-Text NLG with Language Agnostic Delexicalisation · ACL/IJCNLP (1) 2021
Computer vision › Segmentation and scene understanding
instance segmentation
0.212024
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation · CVPR 2024
Natural language and speech › Language models and text generation
natural language understanding
0.212022
Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU · EMNLP 2022
Natural language and speech › Information extraction and text analysis › relation extraction
definition extraction
0.112009
Finding Short Definitions of Terms on Web Pages · EMNLP 2009
Information retrieval
web search
0.112009
Finding Short Definitions of Terms on Web Pages · EMNLP 2009

Methods — techniques the papers use, named apart from their topics

instance completion · 1.5image re-assembly · 1.5image decomposition · 1.5temporally contiguous retrieval · 0.9similarity-based retrieval · 0.9mixture of attention · 0.9large language model · 0.9graph-theoretic boundary refinement · 0.9bayesian surprise · 0.9difficulty metrics · 0.6
YearPublicationVenuePosition
2025 Human-inspired Episodic Memory for Infinite Context LLMs
abstract
Large language models (LLMs) have shown remarkable capabilities, but still struggle with processing extensive contexts, limiting their ability to maintain coherence and accuracy over long sequences. In contrast, the human brain excels at organising and retrieving episodic experiences across vast temporal scales, spanning a lifetime. In this work, we introduce EM-LLM, a novel approach that integrates key aspects of human episodic memory and event cognition into LLMs with no fine-tuning, enabling them to handle practically infinite context lengths while maintaining computational efficiency. EM-LLM organises sequences of tokens into coherent episodic events using a combination of Bayesian surprise and graph-theoretic boundary refinement in an online fashion. When needed, these events are retrieved through a two-stage memory process, combining similarity-based and temporally contiguous retrieval for efficient, human-inspired access to relevant information. Experiments on the LongBench and $\infty$-Bench benchmarks demonstrate EM-LLM's superior performance, consistently outperforming the state-of-the-art retrieval model InfLLM across various baseline LLMs. In addition, EM-LLM outperforms its popular counterpart, RAG, in a wide range of tasks, while requiring similar resources. Notably, EM-LLM's performance even surpasses full-context models in most tasks, while successfully performing retrieval across 10 million tokens -- a scale computationally infeasible for such models. Finally, our analysis reveals strong correlations between EM-LLM's event segmentation and human-perceived events, suggesting parallels between this artificial system and its biological counterpart, thereby offering a novel computational framework for exploring human memory mechanisms.
Zafeirios Fountas, Martin Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang 0012
ICLR5
2025 Mixture of Attentions For Speculative Decoding
abstract
The growth in the number of parameters of Large Language Models (LLMs) has led to a significant surge in computational requirements, making them challenging and costly to deploy. Speculative decoding (SD) leverages smaller models to efficiently propose future tokens, which are then verified by the LLM in parallel. Small models that utilise activations from the LLM currently achieve the fastest decoding speeds. However, we identify several limitations of SD models including the lack of on-policyness during training and partial observability. To address these shortcomings, we propose a more grounded architecture for small models by introducing a Mixture of Attentions for SD. Our novel architecture can be applied in two scenarios: a conventional single device deployment and a novel client-server deployment where the small model is hosted on a consumer device and the LLM on a server. In a single-device scenario, we demonstrate state-of-the-art speedups improving EAGLE-2 by 9.5% and its acceptance length by 25%. In a client-server setting, our experiments demonstrate: 1) state-of-the-art latencies with minimal calls to the server for different network conditions, and 2) in the event of a complete disconnection, our approach can maintain higher accuracy compared to other SD methods and demonstrates advantages over API calls to LLMs, which would otherwise be unable to continue the generation process.
Matthieu Zimmer, Milan Gritta, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang 0012
ICLR3
2024 MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
abstract
Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout conditioning, or image editing techniques which often require hand drawn masks. Nonetheless, pre-existing works struggle to take advantage of the natural instance-level compositionality of scenes due to the typically flat nature of rasterized RGB output images. Towards adressing this challenge, we introduce MuLAn: a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multi-layer, instance-wise RGBA decompositions, and over 100K instance images. To build MuLAn, we developed a training free pipeline which decomposes a monocular RGB image into a stack of RGBA layers comprising of background and isolated instances. We achieve this through the use of pre-trained general-purpose models, and by developing three modules: image decomposition for instance discovery and extraction, instance completion to reconstruct occluded areas, and image re-assembly. We use our pipeline to create MuLAn-COCO and MuLAn-LAION datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance decompo-sition and occlusion information for high quality images, opening up new avenues for text-to-image generative AI re-search. With this, we aim to encourage the development of novel generation and editing technology, in particular layer-wise solutions. MuLAn data resources are available at https://MuLAn-dataset.github.io/.
Petru-Daniel Tudosiu, Yongxin Yang, Steven McDonagh 0001, Gerasimos Lampouras, Ignacio Iacobacci, Sarah Parisot
CVPR6
2024 Text-to-Code Generation with Modality-relative Pre-training
abstract
Large pre-trained language models have recently been expanded and applied to programming language tasks with great success, often through further pre-training of a strictly-natural language model-where training sequences typically contain both natural and (linearised) programming language.Such approaches effectively map both modalities of the sequence into the same embedding space.However, programming language keywords (e.g."while") often have very strictly defined semantics.As such, transfer learning from their natural language usage may not necessarily be beneficial to their code application and vise versa.Assuming an already pre-trained language model, in this work we investigate how sequence tokens can be adapted and represented differently, depending on which modality they belong to, and to the ultimate benefit of the downstream task.We experiment with separating embedding spaces between modalities during further model pretraining with modality-relative training objectives.We focus on text-to-code generation and observe consistent improvements across two backbone models and two test sets, measuring pass@k and a novel incremental variation.1
Fenia Christopoulou, Guchun Zhang, Gerasimos Lampouras
EACL (1)3
2024 HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
abstract
Milan Gritta, Gerasimos Lampouras, Ignacio Iacobacci. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Milan Gritta, Gerasimos Lampouras, Ignacio Iacobacci
NAACL-HLT2
2022 Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLU
abstract
Curriculum Learning (CL) is a technique of training models via ranking examples in a typically increasing difficulty trend with the aim of accelerating convergence and improving generalisability.Current approaches for Natural Language Understanding (NLU) tasks use CL to improve in-distribution data performance often via heuristic-oriented or task-agnostic difficulties.In this work, instead, we employ CL for NLU by taking advantage of training dynamics as difficulty metrics, i.e. statistics that measure the behavior of the model at hand on specific task-data instances during training and propose modifications of existing CL schedulers based on these statistics.Differently from existing works, we focus on evaluating models on in-distribution (ID), out-of-distribution (OOD) as well as zero-shot (ZS) cross-lingual transfer datasets.We show across several NLU tasks that CL with training dynamics can result in better performance mostly on zero-shot cross-lingual transfer and OOD settings with improvements up by 8.5% in certain cases.Overall, experiments indicate that training dynamics can lead to better performing models with smoother training compared to other difficulty metrics while being 20% faster on average.In addition, through analysis we shed light on the correlations of task-specific versus task-agnostic metrics 1 .
Fenia Christopoulou, Gerasimos Lampouras, Ignacio Iacobacci
EMNLP2
2021 Generalising Multilingual Concept-to-Text NLG with Language Agnostic Delexicalisation
abstract
Giulio Zhou, Gerasimos Lampouras. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Giulio Zhou, Gerasimos Lampouras
ACL/IJCNLP (1)2
2021 Conversation Graph: Data Augmentation, Training and Evaluation for Non-Deterministic Dialogue Management
abstract
Task-oriented dialogue systems typically rely on large amounts of high-quality training data or require complex handcrafted rules. However, existing datasets are often limited in size con- sidering the complexity of the dialogues. Additionally, conventional training signal in- ference is not suitable for non-deterministic agent behavior, namely, considering multiple actions as valid in identical dialogue states. We propose the Conversation Graph (ConvGraph), a graph-based representation of dialogues that can be exploited for data augmentation, multi- reference training and evaluation of non- deterministic agents. ConvGraph generates novel dialogue paths to augment data volume and diversity. Intrinsic and extrinsic evaluation across three datasets shows that data augmentation and/or multi-reference training with ConvGraph can improve dialogue success rates by up to 6.4%.
Milan Gritta, Gerasimos Lampouras, Ignacio Iacobacci
Trans. Assoc. Comput. Linguistics2
2016 Imitation learning for language generation from unaligned data
abstract
Natural language generation (NLG) is the task of generating natural language from a meaning representation. Current rule-based approaches require domain-specific and manually constructed linguistic resources, while most machine-learning based approaches rely on aligned training data and/or phrase templates. The latter are needed to restrict the search space for the structured prediction task defined by the unaligned datasets. In this work we propose the use of imitation learning for structured prediction which learns an incremental model that handles the large search space by avoiding explicit enumeration of the outputs. We focus on the Locally Optimal Learning to Search framework which allows us to train against non-decomposable loss functions such as the BLEU or ROUGE scores while not assuming gold standard alignments. We evaluate our approach on three datasets using both automatic measures and human judgements and achieve results comparable to the state-of-the-art approaches developed for each of them.
Gerasimos Lampouras, Andreas Vlachos 0001
COLING1
2013 Generating Natural Language Descriptions from OWL Ontologies: the NaturalOWL System
abstract
We present NaturalOWL, a natural language generation system that produces texts describing individuals or classes of OWL ontologies. Unlike simpler OWL verbalizers, which typically express a single axiom at a time in controlled, often not entirely fluent natural language primarily for the benefit of domain experts, we aim to generate fluent and coherent multi-sentence texts for end-users. With a system like NaturalOWL, one can publish information in OWL on the Web, along with automatically produced corresponding texts in multiple languages, making the information accessible not only to computer programs and domain experts, but also end-users. We discuss the processing stages of NaturalOWL, the optional domain-dependent linguistic resources that the system can use at each stage, and why they are useful. We also present trials showing that when the domain-dependent llinguistic resources are available, NaturalOWL produces significantly better texts compared to a simpler verbalizer, and that the resources can be created with relatively light effort.
Ion Androutsopoulos, Gerasimos Lampouras, Dimitrios Galanis
J. Artif. Intell. Res.2
2012 Extractive Multi-Document Summarization with Integer Linear Programming and Support Vector Regression
Dimitrios Galanis, Gerasimos Lampouras, Ion Androutsopoulos
COLING2
2009 Finding Short Definitions of Terms on Web Pages
Gerasimos Lampouras, Ion Androutsopoulos
EMNLP1