Guy Emerson

dblp:182/2001 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-3136-9682ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 52% Knowledge representation and reasoning · 13% Vision and language · 10%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 77% Information retrieval · 23%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
distributional semantics
2.242024
Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics · ACL (1) 2024
Learning Functional Distributional Semantics with Visual Data · ACL (1) 2022
What are the Goals of Distributional Semantics? · ACL 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › semantic relations
hypernymy detection
0.812024
Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
lexical semantics
0.812024
Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics · ACL (1) 2024
Knowledge graphs
knowledge graph embedding
0.712023
Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics · EMNLP 2023
Knowledge graphs
link prediction
0.712023
Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics · EMNLP 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference
0.412020
Autoencoding Pixies: Amortised Variational Inference with Graph Convolutions for Functional Distributional Semantics · ACL 2020
Machine learning › Generative modeling
variational autoencoder
0.412020
Autoencoding Pixies: Amortised Variational Inference with Graph Convolutions for Functional Distributional Semantics · ACL 2020
Information retrieval › evaluation
benchmark evaluation
0.212023
Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics · EMNLP 2023
Information retrieval
evaluation
0.212023
Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics · EMNLP 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112020
Investigating Cross-Linguistic Adjective Ordering Tendencies with a Latent-Variable Model · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

variational autoencoding · 0.8probing · 0.8link prediction · 0.7knowledge base embedding · 0.7visual grounding · 0.6latent variable model · 0.4graph convolutional network · 0.4corpus analysis · 0.4autoencoder · 0.4
YearPublicationVenuePosition
2024 Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional Semantics
abstract
Functional Distributional Semantics (FDS) models the meaning of words by truthconditional functions.This provides a natural representation for hypernymy but no guarantee that it can be learnt when FDS models are trained on a corpus.In this paper, we probe into FDS models and study the representations learnt, drawing connections between quantifications, the Distributional Inclusion Hypothesis (DIH), and the variational-autoencoding objective of FDS model training.Using synthetic data sets, we reveal that FDS models learn hypernymy on a restricted class of corpus that strictly follows the DIH.We further introduce a training objective that both enables hypernymy learning under the reverse of the DIH and improves hypernymy detection from real corpora.
Chun Hei Lo, Wai Lam, Hong Cheng 0001, Guy Emerson
ACL (1)4
2024 UG-schematic Annotation for Event Nominals: A Case Study in Mandarin Chinese
abstract
Abstract Divergence of languages observed at the surface level is a major challenge encountered by multilingual data representation, especially when typologically distant languages are involved. Drawing inspiration from a formalist Chomskyan perspective towards language universals, Universal Grammar (UG), this article uses deductively pre-defined universals to analyze a multilingually heterogeneous phenomenon, event nominals. In this way, deeper universality of event nominals beneath their huge divergence in different languages is uncovered, which empowers us to break barriers between languages and thus extend insights from some synthetic languages to a non-inflectional language, Mandarin Chinese. Our empirical investigation also demonstrates this UG-inspired schema is effective: With its assistance, the inter-annotator agreement (IAA) for identifying event nominals in Mandarin grows from 88.02% to 94.99%, and automatic detection of event-reading nominalizations on the newly-established data achieves an accuracy of 94.76% and an F1 score of 91.3%, which significantly surpass those achieved on the pre-existing resource by 9.8% and 5.2%, respectively. Our systematic analysis also sheds light on nominal semantic role labeling. By providing a clear definition and classification on arguments of event nominal, the IAA of this task significantly increases from 90.46% to 98.04%.
Wenxi Li, Guy Emerson
Comput. Linguistics3
2023 Are Embedded Potatoes Still Vegetables? On the Limitations of WordNet Embeddings for Lexical Semantics
abstract
Knowledge Base Embedding (KBE) models have been widely used to encode structured information from knowledge bases, including WordNet.However, the existing literature has predominantly focused on link prediction as the evaluation task, often neglecting exploration of the models' semantic capabilities.In this paper, we investigate the potential disconnect between the performance of KBE models of WordNet on link prediction and their ability to encode semantic information, highlighting the limitations of current evaluation protocols.Our findings reveal that some top-performing KBE models on the WN18RR benchmark exhibit subpar results on two semantic tasks and two downstream tasks.These results demonstrate the inadequacy of link prediction benchmarks for evaluating the semantic capabilities of KBE models, suggesting the need for a more targeted assessment approach.
Xuyou Cheng, Michael Sejr Schlichtkrull, Guy Emerson
EMNLP3
2023 Visual Spatial Reasoning
abstract
Abstract Spatial relations are a basic part of human cognition. However, they are expressed in natural language in a variety of ways, and previous work has suggested that current vision-and-language models (VLMs) struggle to capture relational information. In this paper, we present Visual Spatial Reasoning (VSR), a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). While using a seemingly simple annotation format, we show how the dataset includes challenging linguistic phenomena, such as varying reference frames. We demonstrate a large gap between human and model performance: The human ceiling is above 95%, while state-of-the-art models only achieve around 70%. We observe that VLMs’ by-relation performances have little correlation with the number of training examples and the tested models are in general incapable of recognising relations concerning the orientations of objects.1
Fangyu Liu 0001, Guy Emerson, Nigel Collier
Trans. Assoc. Comput. Linguistics2
2022 Learning Functional Distributional Semantics with Visual Data
abstract
Functional Distributional Semantics is a recently proposed framework for learning distributional semantics that provides linguistic interpretability.It models the meaning of a word as a binary classifier rather than a numerical vector.In this work, we propose a method to train a Functional Distributional Semantics model with grounded visual data.We train it on the Visual Genome dataset, which is closer to the kind of data encountered in human language acquisition than a large text corpus.On four external evaluation datasets, our model outperforms previous work on learning semantics from Visual Genome. 1
Yinhong Liu, Guy Emerson
ACL (1)2
2021 Incremental Beam Manipulation for Natural Language Generation
abstract
The performance of natural language generation systems has improved substantially with modern neural networks.At test time they typically employ beam search to avoid locally optimal but globally suboptimal predictions.However, due to model errors, a larger beam size can lead to deteriorating performance according to the evaluation metric.For this reason, it is common to rerank the output of beam search, but this relies on beam search to produce a good set of hypotheses, which limits the potential gains.Other alternatives to beam search require changes to the training of the model, which restricts their applicability compared to beam search.This paper proposes incremental beam manipulation, i.e. reranking the hypotheses in the beam during decoding instead of only at the end.This way, hypotheses that are unlikely to lead to a good final output are discarded, and in their place hypotheses that would have been ignored will be considered instead.Applying incremental beam manipulation leads to an improvement of 1.93 and 5.82 BLEU points over vanilla beam search for the test sets of the E2E and WebNLG challenges respectively.The proposed method also outperformed a strong reranker by 1.04 BLEU points on the E2E challenge, while being on par with it on the WebNLG dataset.
James Hargreaves, Andreas Vlachos 0001, Guy Emerson
EACL3
2020 Autoencoding Pixies: Amortised Variational Inference with Graph Convolutions for Functional Distributional Semantics
abstract
Functional Distributional Semantics provides a linguistically interpretable framework for distributional semantics, by representing the meaning of a word as a function (a binary classifier), instead of a vector. However, the large number of latent variables means that inference is computationally expensive, and training a model is therefore slow to converge. In this paper, I introduce the Pixie Autoencoder, which augments the generative model of Functional Distributional Semantics with a graph-convolutional neural network to perform amortised variational inference. This allows the model to be trained more effectively, achieving better results on two tasks (semantic similarity in context and semantic composition), and outperforming BERT, a large pre-trained language model.
Guy Emerson
ACL1
2020 What are the Goals of Distributional Semantics?
abstract
Distributional semantic models have become a mainstay in NLP, providing useful features for downstream tasks.However, assessing long-term progress requires explicit long-term goals.In this paper, I take a broad linguistic perspective, looking at how well current models can deal with various semantic challenges.Given stark differences between models proposed in different subfields, a broad perspective is needed to see how we could integrate them.I conclude that, while linguistic insights can guide the design of model architectures, future progress will require balancing the often conflicting demands of linguistic expressiveness and computational tractability.
Guy Emerson
ACL1
2020 Investigating Cross-Linguistic Adjective Ordering Tendencies with a Latent-Variable Model
abstract
Across languages, multiple consecutive adjectives modifying a noun (e.g."the big red dog") follow certain unmarked ordering rules.While explanatory accounts have been put forward, much of the work done in this area has relied primarily on the intuitive judgment of native speakers, rather than on corpus data.We present the first purely corpus-driven model of multi-lingual adjective ordering in the form of a latent-variable model that can accurately order adjectives across 24 different languages, even when the training and testing languages are different.We utilize this novel statistical model to provide strong converging evidence for the existence of universal, cross-linguistic, hierarchical adjective ordering tendencies.
Jun Yen Leung, Guy Emerson, Ryan Cotterell
EMNLP (1)2
2016 Resources for building applications with Dependency Minimal Recursion Semantics
Ann A. Copestake, Guy Emerson, Michael Wayne Goodman, Matic Horvat, Ewa Muszynska
LREC2