VLDB 2026 Research / reviewers in the wild / expert
Nina Dethlefs
dblp:63/8341
· DBLP profile ↗
29ranked-venue papers
11as first author
7since 2021 · last 2024
0000-0002-6917-5066ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 11 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 49% Language models and text generation · 35% Question answering and dialogue systems · 9% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
emotion recognition |
0.8 | 1 | 2024 | Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through Paraphrasing · EMNLP 2024 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.8 | 1 | 2024 | Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through Paraphrasing · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.8 | 1 | 2024 | Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through Paraphrasing · EMNLP 2024 |
Natural language and speech › Language models and text generation › text generation
surface realization |
0.2 | 1 | 2013 | Conditional Random Fields for Responsive Surface Realisation using Global Features · ACL (1) 2013 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 1 | 2013 | Conditional Random Fields for Responsive Surface Realisation using Global Features · ACL (1) 2013 |
Natural language and speech › Question answering and dialogue systems
dialogue management |
0.1 | 1 | 2012 | Optimising Incremental Dialogue Decisions Using Information Density for Interactive Systems · EMNLP-CoNLL 2012 |
Natural language and speech › Question answering and dialogue systems › spoken dialogue systems
incremental dialogue systems |
0.1 | 1 | 2012 | Optimising Incremental Dialogue Decisions Using Information Density for Interactive Systems · EMNLP-CoNLL 2012 |
Methods — techniques the papers use, named apart from their topics
large language model prompting · 0.8fine-tuning · 0.8conditional random field · 0.2information density · 0.1decision optimization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through ParaphrasingabstractIn the realm of social media discourse, the integration of slang enriches communication, reflecting the sociocultural identities of users.This study investigates the capability of large language models (LLMs) to paraphrase slang within climate-related tweets from Nigeria and the UK, with a focus on identifying emotional nuances.Using DistilRoBERTa as the baseline model, we observe its limited comprehension of slang.To improve cross-cultural understanding, we gauge the effectiveness of leading LLMs: ChatGPT 4, Gemini, and LLaMA3 in slang paraphrasing.While ChatGPT 4 and Gemini demonstrate comparable effectiveness in slang paraphrasing, LLaMA3 shows less coverage, with all LLMs exhibiting limitations in coverage, especially of Nigerian slang.Our findings underscore the necessity for culturallysensitive LLM development in emotion classification, particularly in non-anglocentric regions. Ifeoluwa Wuraola, Nina Dethlefs, Daniel Marciniak |
EMNLP | 2 |
| 2023 | Multi-channel Convolutional Neural Network for Precise Meme ClassificationabstractThis paper proposes a multi-channel convolutional neural network (MC-CNN) for classifying memes and non-memes. Our architecture is trained and validated on a challenging dataset that includes non-meme formats with textual attributes, which are also circulated online but rarely accounted for in meme classification tasks. Alongside a transfer learning base, two additional channels capture low-level and fundamental features of memes that make them unique from other images with text. We contribute an approach which outperforms previous meme classifiers specifically in live data evaluation, and one that is better able to generalise ‘in the wild’. Our research aims to improve accurate collation of meme content to support continued research in meme content analysis, and meme-related sub-tasks such as harmful content detection. Victoria Sherratt, Kevin Pimbblet, Nina Dethlefs |
ICMR | 3 |
| 2023 | Hierarchical Multiscale Recurrent Neural Networks for Detecting Suicide NotesabstractRecent statistics in suicide prevention show that people are increasingly posting their last words online and with the unprecedented availability of textual data from social media platforms researchers have the opportunity to analyse such data. Furthermore, psychological studies have shown that our state of mind can manifest itself in the linguistic features we use to communicate. In this article, we investigate whether it is possible to automatically identify suicide notes from other types of social media blogs in two document-level classification tasks. The first task aims to identify suicide notes from depressed and blog posts in a balanced dataset, whilst the second experiment looks at how well suicide notes can be classified when there is a vast amount of neutral text data, which makes the task more applicable to real-world scenarios. Furthermore, we perform a linguistic analysis using LIWC (Linguistic Inquiry and Word Count). We present a learning model for modelling long sequences in two experiment series. We achieve an f1-score of88.26percent over the baselines of0.60in experiment 1 and96.1percent over the baseline in experiment 2. Finally, we show through visualisations which features the learning model identifies, these include emotions such as love and personal pronouns. Annika Marie Schoene, Alexander P. Turner, Geeth de Mel, Nina Dethlefs |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Imputation of Partially Observed Water Quality Data Using Self-Attention LSTMabstractPossible sensory failures in monitoring systems result in partially filled data which may lead to erroneous statistical conclusions. This may affect critical systems such as pollutant detectors and anomaly activity detectors. Therefore, imputation becomes necessary to decrease error. This work addresses the missing data problem by experimenting with various methods in the context of a water quality dataset with high miss rates. Compared models chosen make different assumptions about the data which are Generative Adversarial Networks, Multiple Imputation by Chained Equations, Variational Auto-Encoders, and Recurrent Neural Networks. A novel recurrent neural network architecture with self-attention is proposed in which imputation is done in a single pass. The proposed model performs with a lower root mean square error, ranging between 0.012-0.28, in three of the four locations. The self-attention components increase the interpretability of the imputation process at each stage of the network, providing information to domain experts. Onatkut Dagtekin, Nina Dethlefs |
IJCNN | 2 |
| 2022 | RELATE: Generating a linguistically inspired Knowledge Graph for fine-grained emotion classificationabstractSeveral existing resources are available for sentiment analysis (SA) tasks that are used for learning sentiment specific embedding (SSE) representations. These resources are either large, common-sense knowledge graphs (KG) that cover a limited amount of polarities/emotions or they are smaller in size (e.g.: lexicons), which require costly human annotation and cover fine-grained emotions. Therefore using knowledge resources to learn SSE representations is either limited by the low coverage of polarities/emotions or the overall size of a resource. In this paper, we first introduce a new directed KG called ‘RELATE’, which is built to overcome both the issue of low coverage of emotions and the issue of scalability. RELATE is the first KG of its size to cover Ekman’s six basic emotions that are directed towards entities. It is based on linguistic rules to incorporate the benefit of semantics without relying on costly human annotation. The performance of ‘RELATE’ is evaluated by learning SSE representations using a Graph Convolutional Neural Network (GCN). Annika Marie Schoene, Nina Dethlefs, Sophia Ananiadou |
LREC | 2 |
| 2021 | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li |
SIGDIAL | 8 |
| 2021 | A divide-and-conquer approach to neural natural language generation from structured data
Nina Dethlefs, Annika Marie Schoene, Heriberto Cuayáhuitl |
Neurocomputing | 1 |
| 2020 | Improving the Transparency of Deep Neural Networks using Artificial Epigenetic Molecules
George Lacey, Annika Marie Schoene, Nina Dethlefs, Alexander P. Turner |
IJCCI | 3 |
| 2020 | A Dual Transformer Model for Intelligent Decision Support for Maintenance of Wind TurbinesabstractWind energy is one of the fastest-growing sustainable energy sources in the world but relies crucially on efficient and effective operations and maintenance to generate sufficient amounts of energy and reduce downtime of wind turbines and associated costs. Machine learning has been applied to fault prediction in wind turbines, but these predictions have not been supported with suggestions on how to avert and fix faults. We present a data-to-text generation system utilising transformers for generating corrective maintenance strategies for faults using SCADA data capturing the operational status of turbines. We achieve this in two stages: a first stage identifies faults based on SCADA input features and their relevance. A second stage performs content selection for the language generation task and creates maintenance strategies based on phrase-based natural language templates. Experiments show that our dual transformer model achieves an accuracy of up to 96.75% for alarm prediction and up to 75.35% for its choice of maintenance strategies during content-selection. A qualitative analysis shows that our generated maintenance strategies are promising. We make our human- authored maintenance templates publicly available, and include a brief video explaining our approach. Joyjit Chatterjee, Nina Dethlefs |
IJCNN | 2 |
| 2019 | Modularity Within Artificial Gene Regulatory Networks
George Lacey, Annika Marie Schoene, Nina Dethlefs, Alexander P. Turner |
CEC | 3 |
| 2017 | Deep Text Generation - Using Hierarchical Decomposition to Mitigate the Effect of Rare Data Points
Nina Dethlefs, Alexander P. Turner |
LDK | 1 |
| 2016 | Information density and overlap in spoken dialogue
Nina Dethlefs, Helen Hastie, Heriberto Cuayáhuitl, Yanchao Yu, Verena Rieser, Oliver Lemon |
Comput. Speech Lang. | 1 |
| 2015 | Hierarchical reinforcement learning for situated natural language generationabstractAbstract Natural Language Generation systems in interactive settings often face a multitude of choices, given that the communicative effect of each utterance they generate depends crucially on the interplay between its physical circumstances, addressee and interaction history. This is particularly true in interactive and situated settings. In this paper we present a novel approach forsituated Natural Language Generationin dialogue that is based onhierarchical reinforcement learningand learns the best utterance for a context by optimisation through trial and error. The model is trained from human–human corpus data and learns particularly to balance the trade-off betweenefficiencyanddetailin giving instructions: the user needs to be given sufficient information to execute their task, but without exceeding their cognitive load. We present results from simulation and a task-based human evaluation study comparing two different versions of hierarchical reinforcement learning: One operates using a hierarchy of policies with a large state space and local knowledge, and the other additionally shares knowledge across generation subtasks to enhance performance. Results show that sharing knowledge across subtasks achieves better performance than learning in isolation, leading to smoother and more successful interactions that are better perceived by human users. Nina Dethlefs, Heriberto Cuayáhuitl |
Nat. Lang. Eng. | 1 |
| 2014 | Cluster-based Prediction of User Ratings for Stylistic Surface RealisationabstractSurface realisations typically depend on their target style and audience.A challenge in estimating a stylistic realiser from data is that humans vary significantly in their subjective perceptions of linguistic forms and styles, leading to almost no correlation between ratings of the same utterance.We address this problem in two steps.First, we estimate a mapping function between the linguistic features of a corpus of utterances and their human style ratings.Users are partitioned into clusters based on the similarity of their ratings, so that ratings for new utterances can be estimated, even for new, unknown users.In a second step, the estimated model is used to re-rank the outputs of a number of surface realisers to produce stylistically adaptive output.Results confirm that the generated styles are recognisable to human judges and that predictive models based on clusters of users lead to better rating predictions than models based on an average population of users. Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, Oliver Lemon |
EACL | 1 |
| 2014 | A Semi-supervised Clustering Approach for Semantic Slot LabellingabstractWork on training semantic slot labellers for use in Natural Language Processing applications has typically either relied on large amounts of labelled input data, or has assumed entirely unlabelled inputs. The former technique tends to be costly to apply, while the latter is often not as accurate as its supervised counterpart. Here, we present a semi-supervised learning approach that automatically labels the semantic slots in a set of training data and aims to strike a balance between the dependence on labelled data and prediction accuracy. The essence of our algorithm is to cluster clauses based on a similarity function that combines lexical and semantic information. We present experiments that compare different similarity functions for both our semi-supervised setting and a fully unsupervised baseline. While semi-supervised learning expectedly outperforms unsupervised learning, our results show that (1) this effect can be observed based on very few training data instances and that increasing the size of the training data does not lead to better performance, and (2) that lexical and semantic information contribute differently in different domains so that clustering based on both types of information offers the best generalisation. Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie |
ICMLA | 2 |
| 2014 | The PARLANCE mobile application for interactive search in English and MandarinabstractHelen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu |
SIGDIAL Conference | 7 |
| 2014 | Training a statistical surface realiser from automatic slot labellingabstractTraining a statistical surface realiser typically relies on labelled training data or parallel data sets, such as corpora of paraphrases. The procedure for obtaining such data for new domains is not only time-consuming, but it also restricts the incorporation of new semantic slots during an interaction, i.e. using an online learning scenario for automatically extended domains. Here, we present an alternative approach to statistical surface realisation from unlabelled data through automatic semantic slot labelling. The essence of our algorithm is to cluster clauses based on a similarity function that combines lexical and semantic information. Annotations need to be reliable enough to be utilised within a spoken dialogue system. We compare different similarity functions and evaluate our surface realiser—trained from unlabelled data—in a human rating study. Results confirm that a surface realiser trained from automatic slot labels can lead to outputs of comparable quality to outputs trained from human-labelled inputs. Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie, Xingkun Liu |
SLT | 2 |
| 2014 | Introduction to the Special Issue on Machine Learning for Multiple Modalities in Interactive Systems and RobotsabstractThis special issue highlights research articles that apply machine learning to robots and other systems that interact with users through more than one modality, such as speech, gestures, and vision. For example, a robot may coordinate its speech with its actions, taking into account (audio-)visual feedback during their execution. Machine learning provides interactive systems with opportunities to improve performance not only of individual components but also of the system as a whole. However, machine learning methods that encompass multiple modalities of an interactive system are still relatively hard to find. The articles in this special issue represent examples that contribute to filling this gap. Heriberto Cuayáhuitl, Lutz Frommberger, Nina Dethlefs, Antoine Raux, Matthew Marge, Hendrik Zender |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2014 | Nonstrict Hierarchical Reinforcement Learning for Interactive Systems and RobotsabstractConversational systems and robots that use reinforcement learning for policy optimization in large domains often face the problem of limited scalability. This problem has been addressed either by using function approximation techniques that estimate the approximate true value function of a policy or by using a hierarchical decomposition of a learning task into subtasks. We present a novel approach for dialogue policy optimization that combines the benefits of both hierarchical control and function approximation and that allows flexible transitions between dialogue subtasks to give human users more control over the dialogue. To this end, each reinforcement learning agent in the hierarchy is extended with a subtask transition function and a dynamic state space to allow flexible switching between subdialogues. In addition, the subtask policies are represented with linear function approximation in order to generalize the decision making to situations unseen in training. Our proposed approach is evaluated in an interactive conversational robot that learns to play quiz games. Experimental results, using simulation and real users, provide evidence that our proposed approach can lead to more flexible (natural) interactions than strict hierarchical control and that it is preferred by human users. Heriberto Cuayáhuitl, Ivana Kruijff-Korbayová, Nina Dethlefs |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2013 | Conditional Random Fields for Responsive Surface Realisation using Global Features
Nina Dethlefs, Helen Hastie, Heriberto Cuayáhuitl, Oliver Lemon |
ACL (1) | 1 |
| 2013 | Barge-in effects in Bayesian dialogue act recognition and simulationabstractDialogue act recognition and simulation are traditionally considered separate processes. Here, we argue that both can be fruitfully treated as interleaved processes within the same probabilistic model, leading to a synchronous improvement of performance in both. To demonstrate this, we train multiple Bayes Nets that predict the timing and content of the next user utterance. A specific focus is on providing support for barge-ins. We describe experiments using the Let's Go data that show an improvement in classification accuracy (+5%) in Bayesian dialogue act recognition involving barge-ins using partial context compared to using full context. Our results also indicate that simulated dialogues with user barge-in are more realistic than simulations without barge-in events. Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie, Oliver Lemon |
ASRU | 2 |
| 2013 | Impact of ASR N-Best Information on Bayesian Dialogue Act Recognition
Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie, Oliver Lemon |
SIGDIAL Conference | 2 |
| 2013 | Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay |
SIGDIAL Conference | 5 |
| 2012 | Optimising Incremental Dialogue Decisions Using Information Density for Interactive Systems
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon |
EMNLP-CoNLL | 1 |
| 2012 | Optimising Incremental Generation for Spoken Dialogue Systems: Reducing the Need for Fillers
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon |
INLG | 1 |
| 2012 | Comparing HMMs and Bayesian Networks for Surface Realisation
Nina Dethlefs, Heriberto Cuayáhuitl |
HLT-NAACL | 1 |
| 2011 | Optimizing Situated Dialogue Management in Unknown EnvironmentsabstractWe present a conversational learning agent that helps users navigate through complex and challenging spatial environments. The agent exhibits adaptive behaviour by learning spatiallyaware dialogue actions while the user carries out the navigation task. To this end, we use Hierarchical Reinforcement Learning with relational representations to efficiently optimize dialogue actions tightly-coupled with spatial ones, and Bayesian networks to model the user’s beliefs of the navigation environment. Since these beliefs are continuously changing, we induce the agent’s behaviour in real time. Experimental results, using simulation, are encouraging by showing efficient adaptation to the user’s navigation knowledge, specifically to the generated route and the intermediate locations to negotiate with the user. Index Terms: spoken dialogue systems, situated interaction, reinforcement learning, hierarchical control, Bayesian networks Heriberto Cuayáhuitl, Nina Dethlefs |
INTERSPEECH | 2 |
| 2011 | Optimising Natural Language Generation Decision Making For Situated Dialogue
Nina Dethlefs, Heriberto Cuayáhuitl, Jette Viethen |
SIGDIAL Conference | 1 |
| 2010 | Hierarchical Reinforcement Learning for Adaptive Text Generation
Nina Dethlefs, Heriberto Cuayáhuitl |
INLG | 1 |