Frank Keller

dblp:30/4872 · DBLP profile ↗
← Back
69ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-8242-4362ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 since 2021
YearPublicationVenuePosition
2026 Generating Visual Stories with Grounded and Coreferent Characters
abstract
Abstract Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story’s themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative around specific characters. As a result, the generated stories feel generic, with character mentions being absent, vague, or incorrect. To mitigate these issues, we introduce a new character-centric approach to visual story generation. We present the first model capable of predicting visual stories with consistently grounded and coreferent character mentions. Our model is finetuned on a new dataset which we build on top of the widely used VIST (Huang et al., 2016) benchmark. Specifically, we develop an automated pipeline to enrich VIST with visual and textual character coreference chains. We also propose new evaluation metrics to measure the richness of characters and coreference in stories. Experimental results show that our model generates stories with recurring characters which are consistent and coreferent to larger extent compared to baselines and state-of-the-art systems.1 Our code and dataset are available at https://github.com/iz2late/character-centric-vist.
Mirella Lapata, Frank Keller
Trans. Assoc. Comput. Linguistics3
2025 Predicting Implicit Arguments in Procedural Video Instructions
abstract
Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like verb,what,where/with. Procedural instructions are highly elliptic, for instance, (i) add cucumber to the bowl and (ii) add sliced tomatoes, the second step's where argument is inferred from the context, referring to where the cucumber was placed. Prior SRL benchmarks often miss implicit arguments, leading to incomplete understanding. To address this, we introduce Implicit-VidSRL, a dataset that necessitates inferring implicit and explicit arguments from contextual information in multimodal cooking procedures. Our proposed dataset benchmarks multimodal models' contextual reasoning, requiring entity tracking through visual changes in recipes. We study recent multimodal LLMs and reveal that they struggle to predict implicit arguments of what and where/with from multi-modal procedural data given the verb. Lastly, we propose iSRL-Qwen2-VL, which achieves a 17% relative improvement in F1-score for what-implicit and a 14.7% for where/with-implicit semantic roles over GPT-4o.
Anil Batra, Laura Sevilla-Lara, Marcus Rohrbach, Frank Keller
ACL (1)4
2025 Path encoding and manner salience in motion event descriptions: the case of Bulgarian and English
Radina Dobreva, Annie Holtz, Alexandra Birch, Frank Keller
CogSci4
2025 End-to-End Long Document Summarization using Gradient Caching
abstract
Abstract Training transformer-based encoder-decoder models for long document summarization poses a significant challenge due to the quadratic memory consumption during training. Several approaches have been proposed to extend the input length at test time, but training with these approaches is still difficult, requiring truncation of input documents and causing a mismatch between training and test conditions. In this work, we propose CachED (Gradient Caching for Encoder-Decoder models), an approach that enables end-to-end training of existing transformer-based encoder-decoder models, using the entire document without truncation. Specifically, we apply non-overlapping sliding windows to input documents, followed by fusion in decoder. During backpropagation, the gradients are cached at the decoder and are passed through the encoder in chunks by re-computing the hidden vectors, similar to gradient checkpointing. In the experiments on long document summarization, we extend BART to CachED BART, processing more than 500K tokens during training and achieving superior performance without using any additional parameters.
Rohit Saxena, Frank Keller
Trans. Assoc. Comput. Linguistics3
2024 Predicting long context effects using surprisal
Georgia-Ann Carter, Frank Keller, Paul Hoffman
CogSci2
2024 Effects of Context on the Use of Descriptive Verbs
Radina Dobreva, Frank Keller, Alexandra Birch
CogSci2
2024 Efficient Pre-training for Localized Instruction Generation of Procedural Videos
Anil Batra, Davide Moltisanti, Laura Sevilla-Lara, Marcus Rohrbach, Frank Keller
ECCV (39)5
2024 Finding the Right Moment: Human-Assisted Trailer Creation via Task Composition
abstract
Movie trailers perform multiple functions: they introduce viewers to the story, convey the mood and artistic style of the film, and encourage audiences to see the movie. These diverse functions make trailer creation a challenging endeavor. In this work, we focus on finding trailer moments in a movie, i.e., shots that could be potentially included in a trailer. We decompose this task into two subtasks: narrative structure identification and sentiment prediction. We model movies as graphs, where nodes are shots and edges denote semantic relations between them. We learn these relations using joint contrastive training which distills rich textual information (e.g., characters, actions, situations) from screenplays. An unsupervised algorithm then traverses the graph and selects trailer moments from the movie that human judges prefer to ones selected by competitive supervised approaches. A main advantage of our algorithm is that it uses interpretable criteria, which allows us to deploy it in an interactive tool for trailer creation with a human in the loop. Our tool allows users to select trailer shots in under 30 minutes that are superior to fully automatic methods and comparable to (exclusive) manual selection by experts.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Detecting and Grounding Important Characters in Visual Stories
abstract
Characters are essential to the plot of any story. Establishing the characters before writing a story can improve the clarity of the plot and the overall flow of the narrative. However, previous work on visual storytelling tends to focus on detecting objects in images and discovering relationships between them. In this approach, characters are not distinguished from other objects when they are fed into the generation pipeline. The result is a coherent sequence of events rather than a character-centric story. In order to address this limitation, we introduce the VIST-Character dataset, which provides rich character-centric annotations, including visual and textual co-reference chains and importance ratings for characters. Based on this dataset, we propose two new tasks: important character detection and character grounding in visual stories. For both tasks, we develop simple, unsupervised models based on distributional similarity and pre-trained vision-and-language models. Our new dataset, together with these models, can serve as the foundation for subsequent work on analysing and generating stories from a character-centric perspective.
Frank Keller
AAAI2
2023 Leveraging context for perceptual prediction using word embeddings
Georgia-Ann Carter, Paul Hoffman, Frank Keller
CogSci3
2023 Learning Action Changes by Measuring Verb-Adverb Textual Relationships
abstract
The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut “finely”). We cast this problem as a regression task. We measure textual relationships between verbs and adverbs to generate a regression target representing the action change we aim to learn. We test our approach on a range of datasets and achieve state-of-the-art results on both adverb prediction and antonym classification. Furthermore, we outperform previous work when we lift two commonly assumed conditions: the availability of action labels during testing and the pairing of adverbs as antonyms. Existing datasets for adverb recognition are either noisy, which makes learning difficult, or contain actions whose appearance is not influenced by adverbs, which makes evaluation less reliable. To address this, we collect a new high quality dataset: Adverbs in Recipes (AIR). We focus on instructional recipes videos, curating a set of actions that exhibit meaningful visual changes when performed differently. Videos in AIR are more tightly trimmed and were manually reviewed by multiple annotators to ensure high labelling quality. Results show that models learn better from AIR given its cleaner videos. At the same time, adverb prediction on AIR is challenging, demonstrating that there is considerable room for improvement.
Davide Moltisanti, Frank Keller, Hakan Bilen, Laura Sevilla-Lara
CVPR2
2023 Semi-supervised multimodal coreference resolution in image narrations
abstract
In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image.This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of large annotated training sets.To tackle these challenges, we present a data efficient semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context.Our approach incorporates losses for both labeled and unlabeled data within a crossmodal framework.Our evaluation shows that the proposed approach outperforms strong baselines both quantitatively and qualitatively, for the tasks of coreference resolution and narrative grounding.
Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen
EMNLP3
2023 Who are you referring to? Coreference resolution in image narrations
abstract
Coreference resolution aims to identify words and phrases which refer to the same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrations of visual scenes. First, we introduce a new dataset with annotated coreference chains and their bounding boxes, as most existing image-text datasets only contain short sentences without coreferring expressions or labeled chains. We propose a new technique that learns to identify coref-erence chains using weak supervision, only from image-text pairs and a regularization using prior linguistic knowledge. Our model yields large performance gains over several strong baselines in resolving coreferences. We also show that coreference resolution helps improve grounding narratives in images.
Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen
ICCV3
2022 A Closer Look at Temporal Ordering in the Segmentation of Instructional Videos
Anil Batra, Shreyank N. Gowda, Frank Keller, Laura Sevilla-Lara
BMVC3
2022 Modeling Fixation Behavior in Reading with Character-level Neural Attention
Songpeng Yan, Michael Hahn 0001, Frank Keller
CogSci3
2022 Not All Relations are Equal: Mining Informative Labels for Scene Graph Generation
abstract
Scene graph generation (SGG) aims to capture a wide variety of interactions between pairs of objects, which is essential for full scene understanding. Existing SGG methods trained on the entire set of relations fail to acquire complex reasoning about visual and textual correlations due to various biases in training data. Learning on trivial relations that indicate generic spatial configuration like ‘on’ instead of informative relations such as ‘parked on’ does not enforce this complex reasoning, harming generalization. To address this problem, we propose a novel framework for SGG training that exploits relation labels based on their informativeness. Our model-agnostic training procedure imputes missing informative relations for less informative samples in the training data and trains a SGG model on the imputed labels along with existing annotations. We show that this approach can successfully be used in conjunction with state-of-the-art SGG methods and improves their performance significantly in multiple metrics on the standard Visual Genome benchmark. Furthermore, we obtain considerable improvements for unseen triplets in a more challenging zero-shot setting.
Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen
CVPR3
2022 Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition
Shreyank N. Gowda, Marcus Rohrbach, Frank Keller, Laura Sevilla-Lara
ECCV (31)3
2022 CLASTER: Clustering with Reinforcement Learning for Zero-Shot Action Recognition
Shreyank N. Gowda, Laura Sevilla-Lara, Frank Keller, Marcus Rohrbach
ECCV (20)3
2021 Movie Summarization via Sparse Graph Construction
abstract
We summarize full-length movies by creating shorter videos containing their most informative scenes. We explore the hypothesis that a summary can be created by assembling scenes which are turning points (TPs), i.e., key events in a movie that describe its storyline. We propose a model that identifies TP scenes by building a sparse movie graph that represents relations between scenes and is constructed using multimodal information. According to human judges, the summaries created by our approach are more informative and complete, and receive higher ratings, than the outputs of sequence-based models and general-purpose summarization algorithms. The induced graphs are interpretable, displaying different topology for different movie genres.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
AAAI2
2021 Memory and Knowledge Augmented Language Models for Inferring Salience in Long-Form Stories
abstract
Measuring event salience is essential in the understanding of stories.This paper takes a recent unsupervised method for salience detection derived from Barthes Cardinal Functions and theories of surprise and applies it to longer narrative forms.We improve the standard transformer language model by incorporating an external knowledgebase (derived from Retrieval Augmented Generation) and adding a memory mechanism to enhance performance on longer works.We use a novel approach to derive salience annotation using chapteraligned summaries from the Shmoop corpus for classic literary works.Our evaluation against this data demonstrates that our salience detection model improves performance over and above a non-knowledgebase and memory augmented language model, both of which are crucial to this improvement.
David Wilmot, Frank Keller
EMNLP (1)2
2020 Screenplay Summarization Using Latent Narrative Structure
abstract
Most general-purpose extractive summarization models are trained on news articles, which are short and present all important information upfront.As a result, such models are biased by position and often perform a smart selection of sentences from the beginning of the document.When summarizing long narratives, which have complex structure and present information piecemeal, simple position heuristics are not sufficient.In this paper, we propose to explicitly incorporate the underlying structure of narratives into general unsupervised and supervised extractive summarization models.We formalize narrative structure in terms of key narrative events (turning points) and treat it as latent in order to summarize screenplays (i.e., extract an optimal sequence of scenes).Experimental results on the CSI corpus of TV screenplays, which we augment with scene-level summarization labels, show that latent turning points correlate with important aspects of a CSI episode and improve summarization performance over general extractive algorithms, leading to more complete and diverse summaries.Victim: Mike Kimble, found in a Body Farm.Died 6 hours ago, unknown cause of death.CSI discover cow tissue in Mike's body.Cross-contamination is suggested.Probable cause of death: Mike's house has been set on fire.CSI finds blood: Mike was murdered, fire was a cover up.First suspects: Mike's fiance, Jane and her ex-husband, Russ.CSI finds photos in Mike's house of Jane's daughter, Jodie, posing naked.Mike is now a suspect of abusing Jodie.Russ allows CSI to examine his gun.CSI discovers that the bullet that killed Mike was made of frozen beef that melt inside him.They also find beef in Russ' gun.Russ confesses that he knew that Mike was abusing Jody, so he confronted and killed him.CSI discovers that the naked photos were taken on a boat, which belongs to Russ.CSI discovers that it was Russ who was abusing his daughter based on fluids found in his sleeping bag and later killed Mike who tried to help Jodie.Russ is given bail, since no jury would convict a protective father.Russ receives a mandatory life sentence.
Pinelopi Papalampidi, Frank Keller, Lea Frermann, Mirella Lapata
ACL2
2020 Suspense in Short Stories is Predicted By Uncertainty Reduction over Neural Story Representation
abstract
Suspense is a crucial ingredient of narrative fiction, engaging readers and making stories compelling.While there is a vast theoretical literature on suspense, it is computationally not well understood.We compare two ways for modelling suspense: surprise, a backward-looking measure of how unexpected the current state is given the story so far; and uncertainty reduction, a forward-looking measure of how unexpected the continuation of the story is.Both can be computed either directly over story representations or over their probability distributions.We propose a hierarchical language model that encodes stories and computes surprise and uncertainty reduction.Evaluating against short stories annotated with human suspense judgements, we find that uncertainty reduction over representations is the best predictor, resulting in near human accuracy.We also show that uncertainty reduction can be used to predict suspenseful events in movie synopses.
David Wilmot, Frank Keller
ACL2
2019 Dependency Grammar Induction with a Neural Variational Transition-Based Parser
abstract
Dependency grammar induction is the task of learning dependency syntax without annotated training data. Traditional graph-based models with global inference achieve state-ofthe-art results on this task but they require O(n3) run time. Transition-based models enable faster inference with O(n) time complexity, but their performance still lags behind. In this work, we propose a neural transition-based parser for dependency grammar induction, whose inference procedure utilizes rich neural features with O(n) time complexity. We train the parser with an integration of variational inference, posterior regularization and variance reduction techniques. The resulting framework outperforms previous unsupervised transition-based dependency parsers and achieves performance comparable to graph-based models, both on the English Penn Treebank and on the Universal Dependency Treebank. In an empirical comparison, we show that our approach substantially increases parsing speed over graphbased models.
Bowen Li 0002, Jianpeng Cheng 0001, Yang Liu 0124, Frank Keller
AAAI4
2019 An Imitation Learning Approach to Unsupervised Parsing
abstract
Recently, there has been an increasing interest in unsupervised parsers that optimize semantically oriented objectives, typically using reinforcement learning.Unfortunately, the learned trees often do not match actual syntax trees well.Shen et al. (2018) propose a structured attention mechanism for language modeling (PRPN), which induces better syntactic structures but relies on ad hoc heuristics.Also, their model lacks interpretability as it is not grounded in parsing actions.In our work, we propose an imitation learning approach to unsupervised parsing, where we transfer the syntactic knowledge induced by the PRPN to a Tree-LSTM model with discrete parsing actions.Its policy is then refined by Gumbel-Softmax training towards a semantically oriented objective.We evaluate our approach on the All Natural Language Inference dataset and show that it achieves a new state of the art in terms of parsing F -score, outperforming our base models, including the PRPN. 1
Bowen Li 0002, Lili Mou, Frank Keller
ACL (1)3
2019 Character-based Surprisal as a Model of Reading Difficulty in the Presence of Errors
Michael Hahn 0001, Frank Keller, Yonatan Bisk, Yonatan Belinkov
CogSci2
2019 Inattentional Blindness in Visual Search
Matt Rounds, Christopher G. Lucas, Frank Keller
CogSci3
2019 Movie Plot Analysis via Turning Point Identification
abstract
Pinelopi Papalampidi, Frank Keller, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pinelopi Papalampidi, Frank Keller, Mirella Lapata
EMNLP/IJCNLP (1)2
2019 Disambiguating Visual Verbs
abstract
In this article, we introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image. Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration. We introduce a new dataset, which we call VerSe (short for Verb Sense) that augments existing multimodal datasets (COCO and TUHOI) with verb and sense labels. We explore supervised and unsupervised models for the sense disambiguation task using textual, visual, and multimodal embeddings. We also consider a scenario in which we must detect the verb depicted in an image prior to predicting its sense (i.e., there is no verbal information associated with the image). We find that textual embeddings perform well when gold-standard annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images. VerSe is publicly available at https://github.com/spandanagella/verse.
Spandana Gella, Frank Keller, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Training Object Class Detectors with Click Supervision
abstract
Training object class detectors typically requires a large set of images with objects annotated by bounding boxes. However, manually drawing bounding boxes is very time consuming. In this paper we greatly reduce annotation time by proposing center-click annotations: we ask annotators to click on the center of an imaginary bounding box which tightly encloses the object instance. We then incorporate these clicks into existing Multiple Instance Learning techniques for weakly supervised object localization, to jointly localize object bounding boxes over all training images. Extensive experiments on PASCAL VOC 2007 and MS COCO show that: (1) our scheme delivers high-quality detectors, performing substantially better than those produced by weakly supervised techniques, with a modest extra annotation effort, (2) these detectors in fact perform in a range close to those trained from manually drawn bounding boxes, (3) as the center-click task is very fast, our scheme reduces total annotation time by 9x to 18x.
Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller, Vittorio Ferrari
CVPR3
2017 Image Pivoting for Learning Multilingual Multimodal Representations
abstract
In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding.Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages.We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities.We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English.In both cases we achieve state-of-the-art performance.
Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata
EMNLP3
2017 Extreme Clicking for Efficient Object Annotation
abstract
Manually annotating object bounding boxes is central to building computer vision datasets, and it is very time consuming (annotating ILSVRC [53] took 35s for one high-quality box [62]). It involves clicking on imaginary comers of a tight box around the object. This is difficult as these comers are often outside the actual object and several adjustments are required to obtain a tight box. We propose extreme clicking instead: we ask the annotator to click on four physical points on the object: the top, bottom, left- and right-most points. This task is more natural and these points are easy to find. We crowd-source extreme point annotations for PASCAL VOC 2007 and 2012 and show that (1) annotation time is only 7s per box, 5 × faster than the traditional way of drawing boxes [62]: (2) the quality of the boxes is as good as the original ground-truth drawn the traditional way: (3) detectors trained on our annotations are as accurate as those trained on the original ground-truth. Moreover, our extreme clicking strategy not only yields box coordinates, but also four accurate boundary points. We show (4) how to incorporate them into GrabCut to obtain more accurate segmentations than those delivered when initializing it from bounding boxes: (5) semantic segmentations models trained on these segmentations outperform those trained on segmentations derived from bounding boxes.
Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller, Vittorio Ferrari
ICCV3
2017 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures (Extended Abstract)
abstract
Automatic image description generation is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the known approaches based on how they conceptualise this problem and provide a review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image-text datasets and the evaluation measures that have been developed to assess the quality of machine-generated descriptions. Finally we explore future directions in the area of automatic image description.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
IJCAI7
2016 Cross-lingual Transfer of Correlations between Parts of Speech and Gaze Features
abstract
Several recent studies have shown that eye movements during reading provide information about grammatical and syntactic processing, which can assist the induction of NLP models. All these studies have been limited to English, however. This study shows that gaze and part of speech (PoS) correlations largely transfer across English and French. This means that we can replicate previous studies on gaze-based PoS tagging for French, but also that we can use English gaze data to assist the induction of French NLP models.
Maria Barrett, Frank Keller, Anders Søgaard
COLING2
2016 We Don't Need No Bounding-Boxes: Training Object Class Detectors Using Only Human Verification
abstract
Training object class detectors typically requires a large set of images in which objects are annotated by boundingboxes. However, manually drawing bounding-boxes is very time consuming. We propose a new scheme for training object detectors which only requires annotators to verify bounding-boxes produced automatically by the learning algorithm. Our scheme iterates between re-training the detector, re-localizing objects in the training images, and human verification. We use the verification signal both to improve re-training and to reduce the search space for re-localisation, which makes these steps different to what is normally done in a weakly supervised setting. Extensive experiments on PASCAL VOC 2007 show that (1) using human verification to update detectors and reduce the search space leads to the rapid production of high-quality bounding-box annotations, (2) our scheme delivers detectors performing almost as good as those trained in a fully supervised setting, without ever drawing any bounding-box, (3) as the verification task is very quick, our scheme substantially reduces total annotation time by a factor 6×-9×.
Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller, Vittorio Ferrari
CVPR3
2016 Modeling Human Reading with Neural Attention
abstract
When humans read text, they fixate some words and skip others. However, there have been few attempts to explain skipping behavior with computational models, as most existing work has focused on predicting reading times (e.g., using surprisal). In this paper, we propose a novel approach that models both skipping and reading, using an unsupervised architecture that combines a neural attention with autoencoding, trained on raw text using reinforcement learning. Our model explains human reading behavior as a tradeoff between precision of language understanding (encoding the input accurately) and economy of attention (fixating as few words as possible). We evaluate the model on the Dundee eye-tracking corpus, showing that it accurately predicts skipping behavior and reading times, is competitive with surprisal, and captures known qualitative features of human reading.
Michael Hahn 0001, Frank Keller
EMNLP2
2016 Unsupervised Visual Sense Disambiguation for Verbs using Multimodal Embeddings
abstract
We introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image.Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration.We introduce VerSe, a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.We propose an unsupervised algorithm based on Lesk which performs visual sense disambiguation using textual, visual, or multimodal embeddings.We find that textual embeddings perform well when goldstandard textual annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images.We also verify our findings by using the textual and multimodal embeddings as features in a supervised setting and analyse the performance of visual sense disambiguation task.VerSe is made publicly available and can be downloaded at: https://github.com/spandanagella/verse.
Spandana Gella, Mirella Lapata, Frank Keller
HLT-NAACL3
2016 Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
abstract
Automatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the existing approaches based on how they conceptualize this problem, viz., models that cast description as either generation problem or as a retrieval problem over a visual or multimodal representational space. We provide a detailed review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image datasets and the evaluation measures that have been developed to assess the quality of machine-generated image descriptions. Finally we extrapolate future directions in the area of automatic image description generation.
Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank
J. Artif. Intell. Res.7
2015 Semantic Role Labeling Improves Incremental Parsing
abstract
Ioannis Konstas, Frank Keller. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ioannis Konstas, Frank Keller
ACL (1)2
2015 Cumulative Contextual Facilitation in Word Activation and Processing: Evidence from Distributional Modelling
Diego Frassinelli, Frank Keller
CogSci2
2014 Query-by-Example Image Retrieval using Visual Dependency Representations
Desmond Elliott, Victor Lavrenko, Frank Keller
COLING3
2014 Training Object Class Detectors from Eye Tracking Data
Dim P. Papadopoulos, Alasdair D. F. Clarke, Frank Keller, Vittorio Ferrari
ECCV (5)3
2014 Incremental Semantic Role Labeling with Tree Adjoining Grammar
abstract
We introduce the task of incremental semantic role labeling (iSRL), in which semantic roles are assigned to incomplete input (sentence prefixes).iSRL is the semantic equivalent of incremental parsing, and is useful for language modeling, sentence completion, machine translation, and psycholinguistic modeling.We propose an iSRL system that combines an incremental TAG parser with a semantically enriched lexicon, a role propagation algorithm, and a cascade of classifiers.Our approach achieves an SRL Fscore of 78.38% on the standard CoNLL 2009 dataset.It substantially outperforms a strong baseline that combines gold-standard syntactic dependencies with heuristic role assignment, as well as a baseline based on Nivre's incremental dependency parser.
Ioannis Konstas, Frank Keller, Vera Demberg, Mirella Lapata
EMNLP2
2013 Object-based Saliency as a Predictor of Attention in Visual Tasks
Michal Dziemianko, Alasdair D. F. Clarke, Frank Keller
CogSci3
2013 The Effect of Incremental Context on Conceptual Processing: Evidence from Visual World and Reading Experiments
Diego Frassinelli, Frank Keller, Christoph Scheepers
CogSci2
2013 Image Description using Visual Dependency Representations
abstract
Describing the main event of an image involves identifying the objects depicted and predicting the relationships between them.Previous approaches have represented images as unstructured bags of regions, which makes it difficult to accurately predict meaningful relationships between regions.In this paper, we introduce visual dependency representations to capture the relationships between the objects in an image, and hypothesize that this representation can improve image description.We test this hypothesis using a new data set of region-annotated images, associated with visual dependency representations and gold-standard descriptions.We describe two template-based description generation models that operate over visual dependency representations.In an image description task, we find that these models outperform approaches that rely on object proximity or corpus information to generate descriptions on both automatic measures and on human judgements.
Desmond Elliott, Frank Keller
EMNLP2
2013 Exploring the Utility of Joint Morphological and Syntactic Learning from Child-directed Speech
abstract
Children learn various levels of linguistic structure concurrently, yet most existing models of language acquisition deal with only a single level of structure, implicitly assuming a sequential learning process.Developing models that learn multiple levels simultaneously can provide important insights into how these levels might interact synergistically during learning.Here, we present a model that jointly induces syntactic categories and morphological segmentations by combining two well-known models for the individual tasks.We test on child-directed utterances in English and Spanish and compare to single-task baselines.In the morphologically poorer language (English), the model improves morphological segmentation, while in the morphologically richer language (Spanish), it leads to better syntactic categorization.These results provide further evidence that joint learning is useful, but also suggest that the benefits may be different for typologically different languages.
Stella Frank, Frank Keller, Sharon Goldwater
EMNLP2
2013 Incremental, Predictive Parsing with Psycholinguistically Motivated Tree-Adjoining Grammar
abstract
Psycholinguistic research shows that key properties of the human sentence processor are incrementality, connectedness (partial structures contain no unattached nodes), and prediction (upcoming syntactic structure is anticipated). There is currently no broad-coverage parsing model with these properties, however. In this article, we present the first broad-coverage probabilistic parser for PLTAG, a variant of TAG that supports all three requirements. We train our parser on a TAG-transformed version of the Penn Treebank and show that it achieves performance comparable to existing TAG parsers that are incremental but not predictive. We also use our PLTAG model to predict human reading times, demonstrating a better fit on the Dundee eye-tracking corpus than a standard surprisal model.
Vera Demberg, Frank Keller, Alexander Koller
Comput. Linguistics2
2013 Incremental Tree Substitution Grammar for Parsing and Sentence Prediction
abstract
In this paper, we present the first incremental parser for Tree Substitution Grammar (TSG). A TSG allows arbitrarily large syntactic fragments to be combined into complete trees; we show how constraints (including lexicalization) can be imposed on the shape of the TSG fragments to enable incremental processing. We propose an efficient Earley-based algorithm for incremental TSG parsing and report an F-score competitive with other incremental parsers. In addition to whole-sentence F-score, we also evaluate the partial trees that the parser constructs for sentence prefixes; partial trees play an important role in incremental interpretation, language modeling, and psycholinguistics. Unlike existing parsers, our incremental TSG parser can generate partial trees that include predictions about the upcoming words in a sentence. We show that it outperforms an n-gram model in predicting more than one upcoming word.
Federico Sangati, Frank Keller
Trans. Assoc. Comput. Linguistics2
2012 A Bayesian Model of the Effect of Object Context on Visual Attention
Ben Allison, Frank Keller, Moreno I. Coco
CogSci2
2012 The Plausibility of Semantic Properties Generated by a Distributional Model: Evidence from a Visual World Experiment
Diego Frassinelli, Frank Keller
CogSci2
2011 Temporal Dynamics of Scan Patterns in Comprehension and Production
Moreno I. Coco, Frank Keller
CogSci2
2011 Incremental Learning of Target Locations in Visual Search
Michal Dziemianko, Moreno I. Coco, Frank Keller
CogSci3
2011 A Model of Discourse Predictions in Human Sentence Processing
Amit Dubey, Frank Keller, Patrick Sturt
EMNLP2
2010 Syntactic and Semantic Factors in Processing Difficulty: An Integrated Measure
Jeff Mitchell 0001, Mirella Lapata, Vera Demberg, Frank Keller
ACL4
2007 Using Foreign Inclusion Detection to Improve Parsing Performance
Beatrice Alex, Amit Dubey, Frank Keller
EMNLP-CoNLL3
2007 An Information Retrieval Approach to Sense Ranking
Mirella Lapata, Frank Keller
HLT-NAACL2
2006 Integrating Syntactic Priming into an Incremental Probabilistic Parser, with an Application to Psycholinguistic Modeling
abstract
The psycholinguistic literature provides evidence for syntactic priming, i.e., the tendency to repeat structures. This paper describes a method for incorporating priming into an incremental probabilistic parser. Three models are compared, which involve priming of rules between sentences, within sentences, and within coordinate structures. These models simulate the reading time advantage for parallel structures found in human data, and also yield a small increase in overall parsing accuracy.
Amit Dubey, Frank Keller, Patrick Sturt
ACL2
2006 Modelling Semantic Role Pausibility in Human Sentence Processing
Ulrike Padó, Matthew W. Crocker, Frank Keller
EACL3
2006 Priming Effects in Combinatory Categorial Grammar
David Reitter, Julia Hockenmaier, Frank Keller
EMNLP3
2006 Computational Modelling of Structural Priming in Dialogue
David Reitter, Frank Keller, Johanna D. Moore
HLT-NAACL2
2005 Lexicalization in Crosslinguistic Probabilistic Parsing: The Case of French
abstract
This paper presents the first probabilistic parsing results for French, using the recently released French Treebank. We start with an unlexicalized PCFG as a baseline model, which is enriched to the level of Collins' Model 2 by adding lexicalization and subcategorization. The lexicalized sister-head model and a bigram model are also tested, to deal with the flatness of the French Treebank. The bigram model achieves the best performance: 81% constituency F-score and 84% dependency accuracy. All lexicalized models outperform the unlexicalized baseline, consistent with probabilistic parsing results for English, but contrary to results for German, where lexicalization has only a limited effect on parsing performance.
Abhishek Arun, Frank Keller
ACL2
2004 The Entropy Rate Principle as a Predictor of Processing Effort: An Evaluation against Eye-tracking Data
Frank Keller
EMNLP1
2004 The Web as a Baseline: Evaluating the Performance of Unsupervised Web-based Models for a Range of NLP Tasks
Mirella Lapata, Frank Keller
HLT-NAACL2
2003 Probabilistic Parsing for German Using Sister-Head Dependencies
abstract
We present a probabilistic parsing model for German trained on the Negra treebank. We observe that existing lexicalized parsing models using head-head dependencies, while successful for English, fail to outperform an unlexicalized baseline model for German. Learning curves show that this effect is not due to lack of training data. We propose an alternative model that uses sister-head dependencies instead of head-head dependencies. This model outperforms the baseline, achieving a labeled precision and recall of up to 74%. This indicates that sister-head dependencies are more appropriate for treebanks with very flat structures such as Negra. 1
Amit Dubey, Frank Keller
ACL2
2003 Using the Web to Obtain Frequencies for Unseen Bigrams
abstract
This article shows that the Web can be employed to obtain frequencies for bigrams that are unseen in a given corpus. We describe a method for retrieving counts for adjective-noun, noun-noun, and verb-object bigrams from the Web by querying a search engine. We evaluate this method by demonstrating: (a) a high correlation between Web frequencies and corpus frequencies; (b) a reliable correlation between Web frequencies and plausibility judgments; (c) a reliable correlation between Web frequencies and frequencies recreated using class-based smoothing; (d) a good performance of Web frequencies in a pseudo disambiguation task.
Frank Keller, Mirella Lapata
Comput. Linguistics1
2002 Using the Web to Overcome Data Sparseness
abstract
This paper shows that the web can be employed to obtain frequencies for bigrams that are unseen in a given corpus. We describe a method for retrieving counts for adjective-noun, noun-noun, and verb-object bigrams from the web by querying a search engine. We evaluate this method by demonstrating that web frequencies and correlate with frequencies obtained from a carefully edited, balanced corpus. We also perform a task-based evaluation, showing that web frequencies can reliably predict human plausibility judgments.
Frank Keller, Maria Lapata, Olga Ourioupina
EMNLP1
2001 Evaluating Smoothing Algorithms against Plausibility Judgements
abstract
Previous research has shown that the plausibility of an adjective-noun combination is correlated with its corpus co-occurrence frequency. In this paper, we estimate the co-occurrence frequencies of adjective-noun pairs that fail to occur in a 100 million word corpus using smoothing techniques and compare them to human plausibility ratings. Both class-based smoothing and distance-weighted averaging yield frequency estimates that are significant predictors of rated plausibility, which provides independent evidence for the validity of these smoothing techniques.
Maria Lapata, Frank Keller, Scott McDonald 0004
ACL2
1999 Determinants of Adjective-Noun Plausibility
Maria Lapata, Scott McDonald 0004, Frank Keller
EACL3
1995 Towards an Account of Extraposition in HPSG
Frank Keller
EACL1