Tom Hosking

dblp:236/5990 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0002-3867-353XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 58% Reinforcement learning · 18% Probabilistic and Bayesian machine learning · 13%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text generation
paraphrase generation
1.122022
Hierarchical Sketch Induction for Paraphrase Generation · ACL (1) 2022
Factorising Meaning and Form for Intent-Preserving Paraphrasing · ACL/IJCNLP (1) 2021
Machine learning › Reinforcement learning
human feedback
0.812024
Human Feedback is not Gold Standard · ICLR 2024
Natural language and speech › Language models and text generation › text summarization
opinion summarization
0.712023
Attributable and Scalable Opinion Summarization · ACL (1) 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model
0.612022
Hierarchical Sketch Induction for Paraphrase Generation · ACL (1) 2022
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.512021
Factorising Meaning and Form for Intent-Preserving Paraphrasing · ACL/IJCNLP (1) 2021
Interaction techniques and input
annotation
0.212024
Human Feedback is not Gold Standard · ICLR 2024

Methods — techniques the papers use, named apart from their topics

user study · 1.5variational autoencoder · 1.1instruction-tuned models · 0.8instruction-tuned model · 0.8unsupervised learning · 0.7latent space aggregation · 0.7hierarchical discrete latent space · 0.7end-to-end training · 0.6discrete latent variable modeling · 0.6adversarial training · 0.5
YearPublicationVenuePosition
2024 Human Feedback is not Gold Standard
abstract
Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single `preference' score captures. We hypothesise that preference scores are subjective and open to undesirable biases. We critically analyse the use of human feedback for both training and evaluation, to verify whether it fully captures a range of crucial error criteria. We find that while preference scores have fairly good coverage, they under-represent important aspects like factuality. We further hypothesise that both preference scores and error annotation may be affected by confounders, and leverage instruction-tuned models to generate outputs that vary along two possible confounding dimensions: assertiveness and complexity. We find that the assertiveness of an output skews the perceived rate of factuality errors, indicating that human annotations are not a fully reliable evaluation metric or training objective. Finally, we offer preliminary evidence that using human feedback as a training objective disproportionately increases the assertiveness of model outputs. We encourage future work to carefully consider whether preference scores are well aligned with the desired objective.
Tom Hosking, Phil Blunsom, Max Bartolo
ICLR1
2024 Hierarchical Indexing for Retrieval-Augmented Opinion Summarization
abstract
Abstract We propose a method for unsupervised abstractive opinion summarization, that combines the attributability and scalability of extractive approaches with the coherence and fluency of Large Language Models (LLMs). Our method, HIRO, learns an index structure that maps sentences to a path through a semantically organized discrete hierarchy. At inference time, we populate the index and use it to identify and retrieve clusters of sentences containing popular opinions from input reviews. Then, we use a pretrained LLM to generate a readable summary that is grounded in these extracted evidential clusters. The modularity of our approach allows us to evaluate its efficacy at each stage. We show that HIRO learns an encoding space that is more semantically structured than prior work, and generates summaries that are more representative of the opinions in the input reviews. Human evaluation confirms that HIRO generates significantly more coherent, detailed, and accurate summaries.
Tom Hosking, Mirella Lapata
Trans. Assoc. Comput. Linguistics1
2023 Attributable and Scalable Opinion Summarization
abstract
We propose a method for unsupervised opinion summarization that encodes sentences from customer reviews into a hierarchical discrete latent space, then identifies common opinions based on the frequency of their encodings.We are able to generate both abstractive summaries by decoding these frequent encodings, and extractive summaries by selecting the sentences assigned to the same frequent encodings.Our method is attributable, because the model identifies sentences used to generate the summary as part of the summarization process.It scales easily to many hundreds of input reviews, because aggregation is performed in the latent space rather than over long sequences of tokens.We also demonstrate that our appraoch enables a degree of control, generating aspectspecific summaries by restricting the model to parts of the encoding space that correspond to desired aspects (e.g., location or food).Automatic and human evaluation on two datasets from different domains demonstrates that our method generates summaries that are more informative than prior work and better grounded in the input reviews.
Tom Hosking, Mirella Lapata
ACL (1)1
2023 Optimal Transport Posterior Alignment for Cross-lingual Semantic Parsing
abstract
Abstract Cross-lingual semantic parsing transfers parsing capability from a high-resource language (e.g., English) to low-resource languages with scarce training data. Previous work has primarily considered silver-standard data augmentation or zero-shot methods; exploiting few-shot gold data is comparatively unexplored. We propose a new approach to cross-lingual semantic parsing by explicitly minimizing cross-lingual divergence between probabilistic latent variables using Optimal Transport. We demonstrate how this direct guidance improves parsing from natural languages using fewer examples and less training. We evaluate our method on two datasets, MTOP and MultiATIS++SQL, establishing state-of-the-art results under a few-shot cross-lingual regime. Ablation studies further reveal that our method improves performance even without parallel input translations. In addition, we show that our model better captures cross-lingual structure in the latent space to improve semantic representation similarity.1
Tom Sherborne, Tom Hosking, Mirella Lapata
Trans. Assoc. Comput. Linguistics2
2022 Hierarchical Sketch Induction for Paraphrase Generation
abstract
We propose a generative model of paraphrase generation, that encourages syntactic diversity by conditioning on an explicit syntactic sketch.We introduce Hierarchical Refinement Quantized Variational Autoencoders (HRQ-VAE), a method for learning decompositions of dense encodings as a sequence of discrete latent variables that make iterative refinements of increasing granularity.This hierarchy of codes is learned through end-to-end training, and represents fine-to-coarse grained information about the input.We use HRQ-VAE to encode the syntactic form of an input sentence as a path through the hierarchy, allowing us to more easily predict syntactic sketches at test time.Extensive experiments, including a human evaluation, confirm that HRQ-VAE learns a hierarchical representation of the input space, and generates paraphrases of higher quality than previous systems.
Tom Hosking, Mirella Lapata
ACL (1)1
2021 Factorising Meaning and Form for Intent-Preserving Paraphrasing
abstract
Tom Hosking, Mirella Lapata. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tom Hosking, Mirella Lapata
ACL/IJCNLP (1)1