Guanyi Chen

dblp:220/0907 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 11 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal Orchestration
abstract
Modern large language model (LLM) serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill phase and memory-bound decode phase. While current practices attempt to address this by organizing these phases into hybrid batches, such solutions create an inefficient tradeoff that sacrifices either throughput or latency, leaving substantial GPU resources underutilized. For this, we identify two key root causes: 1) the prefill phase suffers from suboptimal compute utilization due to wave quantization and attention bottlenecks, and 2) hybrid batching disproportionately prioritizes latency over throughput, wasting both compute resources and memory bandwidth. To mitigate the issues, we present Bullet, a novel spatial-temporal orchestration system that eliminates these inefficiencies through fine-grained phase coordination. Bullet enables concurrent execution of prefill and decode requests, while dynamically provisioning GPU resources based on real-time performance modeling. By integrating SLO-aware scheduling and adaptive resource allocation, Bullet maximizes GPU utilization without compromising latency targets. Experimental evaluations on real-world workloads demonstrate that Bullet delivers 1.26× average throughput gains (up to 1.55×) over state-of-the-arts, while consistently meeting latency constraints.
Zejia Lin 0001, Hongxin Xu, Guanyi Chen, Zhiguang Chen 0001, Yutong Lu, Xianwei Zhang 0001
ASPLOS (2)3
2026 Enhancing Document-Level Event Causality Identification via Event Semantic Consistency and Causal Hypergraph Modeling
Weiying Chen, Guanyi Chen
KSEM (1)3
2026 Detecting Correlation Efficiently in Very Supercritical Stochastic Block Models: Breaking the Otter's Threshold Barrier
abstract
Consider a pair of sparse correlated stochastic block models \(\mathcal S(n, \tfrac{\lambda}{n}; \epsilon; \mathcal s)\) subsampled from a common parent stochastic block model with two symmetric communities, average degree \(\lambda = O(1)\), divergence parameter \(\epsilon \in (0,1)\) and subsampling probability \(\mathcal s\). For all \(\epsilon \in (0,1)\), we construct a statistic based on the combination of two low-degree polynomials and show that there exists a sufficiently small constant \(\delta = \delta(\epsilon) \gt 0\) and a sufficiently large constant \(\Delta = \Delta(\epsilon, \delta)\) such that when \(\lambda \gt \Delta\) and \(\mathcal s \gt \sqrt{\alpha} - \delta\) where \(\alpha \approx 0.338\) is Otter’s constant, this statistic can distinguish this model and a pair of independent stochastic block models \(\mathcal S(n, \tfrac{\lambda s}{n}, \epsilon)\) with probability \(1 - o(1)\). We also provide an efficient algorithm that approximates this statistic in polynomial time. Our result is the first detection or matching type algorithm that breaks the Otter’s threshold in sparse correlated random graphs. The crux of our statistic’s construction lies in a carefully curated family of multigraphs called decorated trees, which enables effective aggregation of the community signal and graph correlation from the counts of the same decorated tree while suppressing the undesirable correlations among counts of different decorated trees. We believe such construction may be of independent interest.
Guanyi Chen, Shuyang Gong, Zhangsong Li
SODA1
2026 Iterative conceptual query expansion for biomedical information retrieval
Chenlian Zhou, Xinhui Tu, Guanyi Chen, Tingting He 0003, Yu Liu 0167
Expert Syst. Appl.3
2026 Emotional supporters often use multiple strategies in a single turn
Guanyi Chen, Chenlian Zhou
Neurocomputing2
2026 Robust Secure Hybrid Beamforming for Active STAR-RIS-Enabled Integrated Sensing and Communication
Guanyi Chen, Bo Li 0034, Weidang Lu, Gang Wang 0021
IEEE Internet Things J.1
2025 On the Human-level Performance of Visual Question Answering
abstract
Visual7W has been widely used in assessing multiple-choice visual question-answering (VQA) systems. This paper reports on a replicated human experiment on Visual7W with the aim of understanding the human-level performance of VQA. The replication was not entirely successful because human participants performed significantly worse when answering “where”, “when”, and “how” questions in compared to other question types. An error analysis discovered that the failure was a consequence of the non-deterministic distractors in Visual7W. GPT-4V was then evaluated using and was compared to the human-level performance. The results embody that, when evaluating models’ capacity on Visual7W, the performance is not necessarily the higher, the better.
Chenlian Zhou, Guanyi Chen
COLING2
2025 Incorporating Formulaicness in the Automatic Evaluation of Naturalness: A Case Study in Logic-to-Text Generation
abstract
Data-to-text natural language generation (NLG) models may produce outputs that closely mirror the structure of their input. We introduce formulaicness as a measure of the output-to-input structural resemblance, proposing it as an enhancement for reference-less naturalness evaluation. Focusing on logic-to-text generation, we construct a dataset and train a regressor to predict formulaicness scores. We collect human judgments on naturalness and examine how incorporating formulaicness into existing metrics affects alignment with these judgments.
Eduardo Calò, Guanyi Chen, Elias Stengel-Eskin, Albert Gatt, Kees van Deemter
INLG2
2025 Annotating Hallucinations in Question-Answering using Rewriting
abstract
Hallucinations pose a persistent challenge in open-ended question answering (QA). Traditional annotation methods, such as span-labelling, suffer from inconsistency and limited coverage. In this paper, we propose a rewriting-based framework as a new perspective on hallucinations in open-ended QA. We report on an experiment in which annotators are instructed to rewrite LLM-generated answers directly to ensure factual accuracy, with edits automatically recorded. Using the Chinese portion of the Mu-SHROOM dataset, we conduct a controlled rewriting experiment, comparing fact-checking tools (Google vs. GPT-4o), and analysing how tool choice, annotator background, and question openness influence rewriting behaviour. We find that rewriting leads to more hallucinations being identified, with higher inter-annotator agreement, than span-labelling.
Guanyi Chen, Kees van Deemter, Tingting He 0003
INLG2
2025 Analysing Reference Production of Large Language Models
abstract
This study investigates how large language models (LLMs) produce referring expressions (REs) and to what extent their behaviour aligns with human patterns. We evaluate LLM performance in two settings: slot filling, %KvD the conventional task of referring expression generation, where REs are generated within a fixed context, and language generation, where REs are analysed within fully generated texts. Using the WebNLG corpus, we assess how well LLMs capture human variation in reference production and analyse their behaviour by examining the influence of several factors known to affect human reference production, including referential form, syntactic position, recency, and discourse status. Our findings show that (1) task framing significantly affects LLMs’ reference production; (2) while LLMs are sensitive to some of these factors, their referential behaviour consistently diverges from human use; and (3) larger model size does not necessarily yield more human-like variation. These results underscore key limitations in current LLMs’ ability to replicate human referential choices.
Chengzhao Wu, Guanyi Chen, Fahime Same, Tingting He 0003
INLG2
2024 Intrinsic Task-based Evaluation for Referring Expression Generation
abstract
Recently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on WEBNLG, Referring Expressions (REs) generated by the state-of-the-art neural models were not only indistinguishable from the REs in WEBNLG but also from the REs generated by a simple rulebased system.Here, we argue that this limitation could stem from the use of a purely ratings-based human evaluation (which is a common practice in Natural Language Generation).To investigate these issues, we propose an intrinsic task-based evaluation for REG models, in which, in addition to rating the quality of REs, participants were asked to accomplish two meta-level tasks.One of these tasks concerns the referential success of each RE; the other task asks participants to suggest a better alternative for each RE.The outcomes suggest that, in comparison to previous evaluations, the new evaluation protocol assesses the performance of each REG model more comprehensively and makes the participants' ratings more reliable and discriminable.
Guanyi Chen, Fahime Same, Kees van Deemter
ACL (1)1
2024 GPTEval: A Survey on Assessments of ChatGPT and GPT-4
abstract
The emergence of ChatGPT has generated much speculation in the press about its potential to disrupt social and economic systems. Its astonishing language ability has aroused strong curiosity among scholars about its performance in different domains. There have been many studies evaluating the ability of ChatGPT and GPT-4 in different tasks and disciplines. However, a comprehensive review summarizing the collective assessment findings is lacking. The objective of this survey is to thoroughly analyze prior assessments of ChatGPT and GPT-4, focusing on its language and reasoning abilities, scientific knowledge, and ethical considerations. Furthermore, an examination of the existing evaluation methods is conducted, offering several recommendations for future research.
Rui Mao 0010, Guanyi Chen, Xulang Zhang, Frank Guerin, Erik Cambria
LREC/COLING2
2024 Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
abstract
Theoretical linguists have suggested that some languages (e.g., Chinese and Japanese) are “cooler” than other languages based on the observation that the intended meaning of phrases in these languages depends more on their contexts. As a result, many expressions in these languages are shortened, and their meaning is inferred from the context. In this paper, we focus on the omission of the plurality and definiteness markers in Chinese noun phrases (NPs) to investigate the predictability of their intended meaning given the contexts. To this end, we built a corpus of Chinese NPs, each of which is accompanied by its corresponding context, and by labels indicating its singularity/plurality and definiteness/indefiniteness. We carried out corpus assessments and analyses. The results suggest that Chinese speakers indeed drop plurality and definiteness markers very frequently. Building on the corpus, we train a bank of computational models using both classic machine learning models and state-of-the-art pre-trained language models to predict the plurality and definiteness of each NP. We report on the performance of these models and analyse their behaviours.
Guanyi Chen, Kees van Deemter
LREC/COLING2
2023 A Multi-task Learning Model for Gold-two-mention Co-reference Resolution
abstract
The task of resolving repeated objects in natural languages is known as co-reference resolution. It is an important part of modern natural language processing and semantic cognition as these implicit relationships are particularly difficult in natural language understanding in downstream tasks. Mention identification and mention linking are the two sub-tasks in the general co-reference resolution research community. Gold-two-mention style co-reference resolution is a special type of co-reference resolution that focuses on linking the ambiguous pronoun to one of the two candidate antecedents. In this paper, we proposed a joint learning model that learns mention identification and mention linking tasks together, because we find that the learning of mention identification can provide supportive dependent information for the learning of mention linking. As far as we know, we propose the first model that introduces a multi-task learning framework to the gold-two-mention co-reference resolution task. We find that our proposed model outperforms state-of-the-art baselines and a single-task learning model on three gold-two-mention co-reference resolution datasets. By comparing the errors made by either the single-task learning model or the multi-task learning model, our error analysis also yields interesting findings about in which way our multi-task learning model makes fewer resolution errors.
Ruicheng Liu, Guanyi Chen, Rui Mao 0010, Erik Cambria
IJCNN2
2023 Models of reference production: How do they withstand the test of time?
abstract
In recent years, many NLP studies have focused solely on performance improvement.In this work, we focus on the linguistic and scientific aspects of NLP.We use the task of generating referring expressions in context (REG-incontext) as a case study and start our analysis from GREC, a comprehensive set of shared tasks in English that addressed this topic over a decade ago.We ask what the performance of models would be if we assessed them (1) on more realistic datasets, and (2) using more advanced methods.We test the models using different evaluation metrics and feature selection experiments.We conclude that GREC can no longer be regarded as offering a reliable assessment of models' ability to mimic human reference production, because the results are highly impacted by the choice of corpus and evaluation metrics.Our results also suggest that pre-trained language models are less dependent on the choice of corpus than classic Machine Learning models, and therefore make more robust class predictions.
Fahime Same, Guanyi Chen, Kees van Deemter
INLG2
2023 Neural referential form selection: Generalisability and interpretability
abstract
In recent years, a range of Neural Referring Expression Generation (REG) systems have been built and they have often achieved encouraging results. However, these models are often thought to lack transparency and generality. Firstly, it is hard to understand what these neural REG models can learn and to compare their performance with existing linguistic theories. Secondly, it is unclear whether they can generalise to data in different text genres and different languages. To answer these questions, we propose to focus on a sub-task of REG: Referential Form Selection (RFS). We introduce the task of RFS and a series of neural RFS models built on state-of-the-art neural REG models. To address the issue of interpretability, we probe these RFS models using probing classifiers that consider information known to impact the human choice of Referential Forms. To address the issue of generalisability, we assess the performance of RFS models on multiple datasets in multiple genres and two different languages, namely, English and Chinese.
Guanyi Chen, Fahime Same, Kees van Deemter
Comput. Speech Lang.1
2023 Computational Modelling of Quantifier Use: Corpus, Models, and Evaluation
abstract
A prominent strand of work in formal semantics investigates the ways in which human languages quantify the elements of a set, as when we say All A are B, Few A are B, and so on. Building on a growing body of empirical studies that shed light on the meaning and the use of quantifiers, we extend this line of work by computationally modelling how human speakers textually describe complex scenes in which quantitative relations play an important role. To this end, we conduct a series of elicitation experiments in which human speakers were asked to perform a linguistic task that invites the use of quantified expressions. The experiments result in a corpus, called QTUNA, made up of short texts that contain a large variety of quantified expressions. We analyse QTUNA, summarise our findings, and explain how we design computational models of human quantifier use accordingly. Finally, we evaluate these models in accordance with QTUNA.
Guanyi Chen, Kees van Deemter
J. Artif. Intell. Res.1
2022 Non-neural Models Matter: a Re-evaluation of Neural Referring Expression Generation Systems
abstract
In recent years, neural models have often outperformed rule-based and classic Machine Learning approaches in NLG.These classic approaches are now often disregarded, for example when new neural models are evaluated.We argue that they should not be overlooked, since for some tasks, well-designed non-neural approaches achieve better performance than neural ones.In this paper, the task of generating referring expressions in linguistic context is used as an example.We examined two very different English datasets (WEBNLG and WSJ), and evaluated each algorithm using both automatic and human evaluations.Overall, the results of these evaluations suggest that rule-based systems with simple rule sets achieve on-par or better performance on both datasets compared to state-of-the-art neural REG systems.In the case of the more realistic dataset, WSJ, a machine learning-based system with well-designed linguistic features performed best.We hope that our work can encourage researchers to consider non-neural models in future.* Equal contribution.Order determined by swapping the order in Chen et al. (2021).
Fahime Same, Guanyi Chen, Kees van Deemter
ACL (1)2
2022 Recognition for Underground Voids in C-Scans Based on GRU Using Ground Penetrating Radar
abstract
Ground Penetrating Radar, as a non-invasive detection instrument, is widely used for shallow underground environment exploring. However, as the interpretation for underground voids is still manually performed, the process is inefficient. Aiming at the challenges of automatic recognition for underground voids, We proposed a recognition algorithm based on Gated Recurrent Unit(GRU), which can recognize 3D underground voids automatically. First, we perform energy detection on the images to obtain the preprocessed and prescreened data. Next, EHD, HOG, and Log-Gabor filters are used to extract features. As the traditional methods can only be applied to 2D images, preprocessing for C-scans is employed. Finally, the aggregated features are fed into GRU for classification. In the experiment, the method is evaluated on simulation data sets, and obtaining a 90% accuracy, which proved the effectiveness and efficiency of our method.
Pengfei Feng, Guanyi Chen, Jinglong Liu, Zhitao Wen, Haoxiang Tian 0002
IGARSS4
2022 Recognition For Underground Voids in C-SCANS Based On LSTM Using Ground Penetrating Radar
abstract
Ground Penetrating Radar, as a non-invasive detection instrument, is widely used for shallow underground environment exploring. However, as the interpretation for underground voids is still manually performed, the process is inefficient. Aiming at the challenges of automatic recognition for underground voids, we proposed a recognition algorithm based on Long Short-Term Memory(LSTM), which can recognize 3D underground voids automatically. First, we perform energy detection on the images to obtain the preprocessed and prescreened data. Next, EHD, HOG, and Log-Gabor filters are used to extract features. As the traditional methods can only be applied to 2D images, preprocessing for C-scans is employed. Finally, the aggregated features are fed into LSTM for classification. In the experiment, the method is evaluated on simulation data sets, and obtaining a 90% accuracy, which proved the effectiveness and efficiency of our method.
Pengfei Feng, Guanyi Chen, Jinglong Liu, Zhitao Wen, Haoxiang Tian 0002
IGARSS4
2022 MMChat: Multi-Modal Chat Dataset on Social Media
abstract
Incorporating multi-modal contexts in conversation is an important step for developing more engaging dialogue systems. In this work, we explore this direction by introducing MMChat: a large scale Chinese multi-modal dialogue corpus (32.4M raw dialogues and 120.84K filtered dialogues). Unlike previous corpora that are crowd-sourced or collected from fictitious movies, MMChat contains image-grounded dialogues collected from real conversations on social media, in which the sparsity issue is observed. Specifically, image-initiated dialogues in common communications may deviate to some non-image-grounded topics as the conversation proceeds. To better investigate this issue, we manually annotate 100K dialogues from MMChat and further filter the corpus accordingly, which yields MMChat-hf. We develop a benchmark model to address the sparsity issue in dialogue generation tasks by adapting the attention routing mechanism on image features. Experiments demonstrate the usefulness of incorporating image features and the effectiveness in handling the sparsity of image features.
Yinhe Zheng, Guanyi Chen, Jian Sun 0021
LREC2
2021 Subsurface Voids Detection from Limited Ground Penetrating Radar Data Using Generative Adversarial Network and YOLOV5
abstract
Recently, preventing road collapse caused by subsurface voids became an urgent problem that needs to be solved in the urban road safety area. While the most study in the field of subsurface object detection by deep learning method has only focused on the objects that can be acquired GPR B-scan data easily. This paper aims to realize the subsurface voids detection under the condition that lacks verified GPR B-Scan data. In this paper, a GPR B-Scan image augmentation method by SinGAN is proposed and the YOLOv5 object detection algorithm is applied correspondingly to detect subsurface voids from both real collected and generated GPR B-Scan data. The detection results show that the proposed technique realized subsurface voids detection in limited verified GPR B-scan data samples and become an inspiration for similar tasks that lacking training samples.
Guanyi Chen, Gang Wang 0021, Xuerong Luo, Mingjie Ji, Pengfei Feng
IGARSS1
2021 What can Neural Referential Form Selectors Learn?
abstract
Despite achieving encouraging results, neural Referring Expression Generation models are often thought to lack transparency.We probed neural Referential Form Selection (RFS) models to find out to what extent the linguistic features influencing the RE form are learnt and captured by state-of-the-art RFS models.The results of 8 probing tasks show that all the defined features were learnt to some extent.The probing tasks pertaining to referential status and syntactic position exhibited the highest performance.The lowest performance was achieved by the probing models designed to predict discourse structure properties beyond the sentence level.
Guanyi Chen, Fahime Same, Kees van Deemter
INLG1
2021 Using BERT for choosing classifiers in Mandarin
abstract
Choosing the most suitable classifier in a linguistic context is a well-known problem in the production of Mandarin and many other languages.The present paper proposes a solution based on BERT, compares this solution to previous neural and rule-based models, and argues that the BERT model performs particularly well on those difficult cases where the classifier adds information to the text.
Jani Järnfors, Guanyi Chen, Kees van Deemter, Rint Sybesma
INLG2
2021 Affective Decoding for Empathetic Response Generation
abstract
Understanding speaker's feelings and producing appropriate responses with emotion connection is a key communicative skill for empathetic dialogue systems.In this paper, we propose a simple technique called Affective Decoding for empathetic response generation.Our method can effectively incorporate emotion signals during each decoding step, and can additionally be augmented with an auxiliary dual emotion encoder, which learns separate embeddings for the speaker and listener given the emotion base of the dialogue.Extensive empirical studies show that our models are perceived to be more empathetic by human evaluations, in comparison to several strong mainstream methods for empathetic responding.
Chengkun Zheng, Guanyi Chen, Chenghua Lin 0002, Ruizhe Li 0001
INLG2
2021 Highly Efficient Knowledge Graph Embedding Learning with Orthogonal Procrustes Analysis
abstract
Knowledge Graph Embeddings (KGEs) have been intensively explored in recent years due to their promise for a wide range of applications. However, existing studies focus on improving the final model performance without acknowledging the computational cost of the proposed approaches, in terms of execution time and environmental impact. This paper proposes a simple yet effective KGE framework which can reduce the training time and carbon footprint by orders of magnitudes compared with state-of-the-art approaches, while producing competitive performance. We highlight three technical innovations: full batch learning via relational matrices, closed-form Orthogonal Procrustes Analysis for KGEs, and non-negative-sampling training. In addition, as the first KGE method whose entity embeddings also store full relation information, our trained models encode rich semantics and are highly interpretable. Comprehensive experiments and ablation studies involving 13 strong baselines and two standard datasets verify the effectiveness and efficiency of our algorithm.
Xutan Peng, Guanyi Chen, Chenghua Lin 0002, Mark Stevenson 0001
NAACL-HLT2
2020 Improving Variational Autoencoder for Text Modelling with Timestep-Wise Regularisation
abstract
The Variational Autoencoder (VAE) is a popular and powerful model applied to text modelling to generate diverse sentences.However, an issue known as posterior collapse (or KL loss vanishing) happens when the VAE is used in text modelling, where the approximate posterior collapses to the prior, and the model will totally ignore the latent variables and be degraded to a plain language model during text generation.Such an issue is particularly prevalent when RNN-based VAE models are employed for text modelling.In this paper, we propose a simple, generic architecture called Timestep-Wise Regularisation VAE (TWR-VAE), which can effectively avoid posterior collapse and can be applied to any RNN-based VAE models.The effectiveness and versatility of our model are demonstrated in different tasks, including language modelling and dialogue response generation.
Ruizhe Li 0001, Xiao Li 0041, Guanyi Chen, Chenghua Lin 0002
COLING3
2020 DGST: a Dual-Generator Network for Text Style Transfer
abstract
We propose DGST, a novel and simple Dual-Generator network architecture for text Style Transfer.Our model employs two generators only, and does not rely on any discriminators or parallel corpus for training.Both quantitative and qualitative experiments on the Yelp and IMDb datasets show that our model gives competitive performance compared to several strong baselines with more complicated architecture designs.
Xiao Li 0041, Guanyi Chen, Chenghua Lin 0002, Ruizhe Li 0001
EMNLP (1)2
2020 Lessons from Computational Modelling of Reference Production in Mandarin and English
abstract
Referring expression generation (REG) algorithms offer computational models of the production of referring expressions.In earlier work, a corpus of referring expressions (REs) in Mandarin was introduced.In the present paper, we annotate this corpus, evaluate classic REG algorithms on it, and compare the results with earlier results on the evaluation of REG for English referring expressions.Next, we offer an in-depth analysis of the corpus, focusing on issues that arise from the grammar of Mandarin.We discuss shortcomings of previous REG evaluations that came to light during our investigation and we highlight some surprising results.Perhaps most strikingly, we found a much higher proportion of under-specified expressions than previous studies had suggested, not just in Mandarin but in English as well.
Guanyi Chen, Kees van Deemter
INLG1
2020 Listener's Social Identity Matters in Personalised Response Generation
abstract
Personalised response generation enables generating human-like responses by means of assigning the generator a social identity.However, pragmatics theory suggests that human beings adjust the way of speaking based on not only who they are but also whom they are talking to.In other words, when modelling personalised dialogues, it might be favourable if we also take the listener's social identity into consideration.To validate this idea, we use gender as a typical example of a social variable to investigate how the listener's identity influences the language used in Chinese dialogues on social media.Also, we build personalised generators.The experiment results demonstrate that the listener's identity indeed matters in the language use of responses and that the response generator can capture such differences in language use.More interestingly, by additionally modelling the listener's identity, the personalised response generator performs better in its own identity.
Guanyi Chen, Yinhe Zheng, Yupei Du
INLG1
2020 Gradations of Error Severity in Automatic Image Descriptions
abstract
Earlier research has shown that evaluation metrics based on textual similarity (e.g., BLEU, CIDEr, Meteor) do not correlate well with human evaluation scores for automatically generated text.We carried out an experiment with Chinese speakers, where we systematically manipulated image descriptions to contain different kinds of errors.Because our manipulated descriptions form minimal pairs with the reference descriptions, we are able to assess the impact of different kinds of errors on the perceived quality of the descriptions.Our results show that different kinds of errors elicit significantly different evaluation scores, even though all erroneous descriptions differ in only one character from the reference descriptions.Evaluation metrics based solely on textual similarity are unable to capture these differences, which (at least partially) explains their poor correlation with human judgments.Our work provides the foundations for future work, where we aim to understand why different errors are seen as more or less severe.
Emiel van Miltenburg, Wei-Ting Lu, Emiel Krahmer, Albert Gatt, Guanyi Chen, Kees van Deemter
INLG5
2020 Out-of-Domain Detection for Natural Language Understanding in Dialog Systems
abstract
Natural Language Understanding (NLU) is a vital component of dialogue systems, and its ability to detect Out-of-Domain (OOD) inputs is critical in practical applications, since the acceptance of the OOD input that is unsupported by the current system may lead to catastrophic failure. However, most existing OOD detection methods rely heavily on manually labeled OOD samples and cannot take full advantage of unlabeled data. This limits the feasibility of these models in practical applications. In this paper, we propose a novel model to generate high-quality pseudo OOD samples that are akin to IN-Domain (IND) input utterances and thereby improves the performance of OOD detection. To this end, an autoencoder is trained to map an input utterance into a latent code. Moreover, the codes of IND and OOD samples are trained to be indistinguishable by utilizing a generative adversarial network. To provide more supervision signals, an auxiliary classifier is introduced to regularize the generated OOD samples to have indistinguishable intent labels. Experiments show that these pseudo OOD samples generated by our model can be used to effectively improve OOD detection in NLU. Besides, we also demonstrate that the effectiveness of these pseudo OOD data can be further improved by efficiently utilizing unlabeled data.
Yinhe Zheng, Guanyi Chen, Minlie Huang
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 A Dual-Attention Hierarchical Recurrent Neural Network for Dialogue Act Classification
abstract
This is a repository copy of A dual-attention hierarchical recurrent neural network for dialogue act classification.
Ruizhe Li 0001, Chenghua Lin 0002, Matthew Collinson, Xiao Li 0041, Guanyi Chen
CoNLL5
2019 Generating Quantified Descriptions of Abstract Visual Scenes
abstract
Quantified expressions have always taken up a central position in formal theories of meaning and language use.Yet quantified expressions have so far attracted far less attention from the Natural Language Generation community than, for example, referring expressions.In an attempt to start redressing the balance, we investigate a recently developed corpus in which quantified expressions play a crucial role; the corpus is the result of a carefully controlled elicitation experiment, in which human participants were asked to describe visually presented scenes.Informed by an analysis of this corpus, we propose algorithms that produce computer-generated descriptions of a wider class of visual scenes, and we evaluate the descriptions generated by these algorithms in terms of their correctness, completeness, and human-likeness.We discuss what this exercise can teach us about the nature of quantification and about the challenges posed by the generation of quantified expressions.
Guanyi Chen, Kees van Deemter, Chenghua Lin 0002
INLG1
2019 QTUNA: A Corpus for Understanding How Speakers Use Quantification
abstract
A prominent strand of work in formal semantics investigates the ways in which human languages quantify over the elements of a set, as when we say "All A are B", "All except two A are B", "Only a few of the A are B" and so on.Our aim is to build Natural Language Generation algorithms that mimic humans' use of quantified expressions.To inform these algorithms, we conducted on a series of elicitation experiments in which human speakers were asked to perform a linguistic task that invites the use of quantified expressions.We discuss how these experiments were conducted and what corpora they gave rise to.We conduct an informal analysis of the corpora, and offer an initial assessment of the challenges that these corpora pose for Natural Language Generation.The dataset is available at: https: //github.com/a-quei/qtuna.
Guanyi Chen, Kees van Deemter, Silvia Pagliaro, Louk Smalbil, Chenghua Lin 0002
INLG1
2019 A Closer Look at Recent Results of Verb Selection for Data-to-Text NLG
abstract
Automatic natural language generation systems need to use the contextually-appropriate verbs when describing different kinds of facts or events, which has triggered research interest on verb selection for data-to-text generation. In this paper, we discuss a few limitations of the current task settings and the evaluation metrics. We also provide two simple, efficient, interpretable baseline approaches for statistical selection of trend verbs, which give a strong performance on both previously used evaluation metrics and our new evaluation.
Guanyi Chen, Jin-Ge Yao
INLG1
2019 Residual Energy Optimization for MIMO SWIPT Two-Way Relaying System
abstract
In this paper, we investigate the multiple-input- multiple-output (MIMO) two-way amplify-and-forward (AF) relaying system with simultaneous wireless information and power transfer (SWIPT), where two users harvest energy from the relay signal by power splitting (PS) scheme. Under the transmit power constraints at all nodes, the energy optimization problem is formulated which maximizes the total residual energy of two users by optimizing the splitting ratios and the precoding matrixes at the relay and two users, while still guaranteeing sufficient signal-to-noise ratios (SNR) at two user nodes. We divide the complex non-convex objective problem into three subproblems which are then solved through the proposed alternating optimization (AO) scheme. Numerical results are provided to verify the analysis and demonstrate the efficiency of the proposed scheme.
Guanyi Chen, Jinlong Wang 0004, Gang Wang 0021, Yikun Zou, Donglai Zhao
VTC Fall1
2018 SimpleNLG-ZH: a Linguistic Realisation Engine for Mandarin
abstract
We introduce SimpleNLG-ZH, a realisation engine for Mandarin that follows the software design paradigm of SimpleNLG (Gatt and Reiter, 2009).We explain the core grammar (morphology and syntax) and the lexicon of SimpleNLG-ZH, which is very different from English and other languages for which SimpleNLG engines have been built.The system was evaluated by regenerating expressions from a body of test sentences and a corpus of humanauthored expressions.Human evaluation was conducted to estimate the quality of regenerated sentences.
Guanyi Chen, Kees van Deemter, Chenghua Lin 0002
INLG1
2018 Modelling Pro-drop with the Rational Speech Acts Model
abstract
We extend the classic Referring Expressions Generation task by considering zero pronouns in "pro-drop" languages such as Chinese, modelling their use by means of the Bayesian Rational Speech Acts model (Frank and Goodman, 2012).By assuming that highly salient referents are most likely to be referred to by zero pronouns (i.e., pro-drop is more likely for salient referents than the less salient ones), the model offers an attractive explanation of a phenomenon not previously addressed probabilistically.
Guanyi Chen, Kees van Deemter, Chenghua Lin 0002
INLG1