Alistair Willis

dblp:67/4642 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0001-7233-8623ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 5Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 67% Recommender systems · 33%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 91% Empirical software engineering · 9%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › personalized search
group-based personalization
0.212014
Improving search personalisation with dynamic group formation · SIGIR 2014
Information retrieval
personalized search
0.212014
Improving search personalisation with dynamic group formation · SIGIR 2014
Recommender systems
user modeling
0.212014
Improving search personalisation with dynamic group formation · SIGIR 2014
Requirements engineering and software design › requirements quality
ambiguity detection
0.112010
Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010
Requirements engineering and software design › requirements specification
natural language requirements
0.112010
Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010
Requirements engineering and software design
requirements quality
0.112010
Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010
Empirical software engineering
developer studies
0.012010
Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic in computer science › logical foundations
formal semantics
0.011999
Two Accounts of Scope Availability and Semantic Underspecification · ACL 1999
Algorithms and data structures
polynomial-time algorithms
0.011999
Two Accounts of Scope Availability and Semantic Underspecification · ACL 1999

Methods — techniques the papers use, named apart from their topics

machine learning · 0.1heuristics · 0.1underspecification · 0.0formal semantics · 0.0
YearPublicationVenuePosition
2020 Identifying Annotator Bias: A new IRT-based method for bias identification
abstract
A basic step in any annotation effort is the measurement of the Inter Annotator Agreement (IAA).An important factor that can affect the IAA is the presence of annotator bias.In this paper we introduce a new interpretation and application of the Item Response Theory (IRT) to detect annotators' bias.Our interpretation of IRT offers an original bias identification method that can be used to compare annotators' bias and characterise annotation disagreement.Our method can be used to spot outlier annotators, improve annotation guidelines and provide a better picture of the annotation reliability.Additionally, because scales for IAA interpretation are not generally agreed upon, our bias identification method is valuable as a complement to the IAA value which can help with understanding the annotation disagreement.
Jacopo Amidei, Paul Piwek, Alistair Willis
COLING3
2020 Developing Students' Written Communication Skills with Jupyter Notebooks
abstract
Written communication skills are considered to be highly desirable in computing graduates. However, many computing students do not have a background in which these skills have been developed, and the skills are often not well addressed within a computing curriculum. For some multidisciplinary areas, such as data science, the range of potential stakeholders makes the need for communications skills all the greater. As interest in data science increases and the technical skills of the area are in ever higher demand, understanding effective teaching and learning of these interdisciplinary aspects is receiving significant attention by academics, industry and government in an effort to address the digital skills gap. In this paper, we report on the experience of adapting a final year data science module in an undergraduate computing curriculum to help develop the skills needed for writing extended reports. From its inception, the module has used Jupyter notebooks to develop the students' skills in the coding aspects of the module. However, over several presentations, we have investigated how the cell-based structure of the notebooks can be exploited to improve the students' understanding of how to structure a report on a data investigation. We have increasingly designed the assessment for the module to take advantage of the learning affordances of Jupyter notebooks to support both raw data analysis and effective report writing. We reflect on the lessons learned from these changes to the assessment model, and the students' responses to the changes.
Alistair Willis, Patricia Charlton, Tony Hirst
SIGCSE1
2019 Agreement is overrated: A plea for correlation to assess human evaluation reliability
abstract
Inter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability.According to existing scales of IAA interpretationsee, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) -most data collected for NLG evaluation fail the reliability test.We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study).Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. ( 2018b), such phenomena can be explained in terms of irreducible human language variability.Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability.Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability.This is illustrated using the three case studies.
Jacopo Amidei, Paul Piwek, Alistair Willis
INLG3
2019 The use of rating and Likert scales in Natural Language Generation human evaluation tasks: A review and some recommendations
abstract
Rating and Likert scales are widely used in evaluation experiments to measure the quality of Natural Language Generation (NLG) systems.We review the use of rating and Likert scales for NLG evaluation tasks published in NLG specialized conferences over the last ten years (135 papers in total).Our analysis brings to light a number of deviations from good practice in their use.We conclude with some recommendations about the use of such scales.Our aim is to encourage the appropriate use of evaluation methodologies in the NLG community.
Jacopo Amidei, Paul Piwek, Alistair Willis
INLG3
2018 Rethinking the Agreement in Human Evaluation Tasks
abstract
Human evaluations are broadly thought to be more valuable the higher the inter-annotator agreement. In this paper we examine this idea. We will describe our experiments and analysis within the area of Automatic Question Generation. Our experiments show how annotators diverge in language annotation tasks due to a range of ineliminable factors. For this reason, we believe that annotation schemes for natural language generation tasks that are aimed at evaluating language quality need to be treated with great care. In particular, an unchecked focus on reduction of disagreement among annotators runs the danger of creating generation goals that reward output that is more distant from, rather than closer to, natural human-like language. We conclude the paper by suggesting a new approach to the use of the agreement metrics in natural language generation evaluation tasks.
Jacopo Amidei, Paul Piwek, Alistair Willis
COLING3
2018 Evaluation methodologies in Automatic Question Generation 2013-2018
abstract
In the last few years Automatic Question Generation (AQG) has attracted increasing interest.In this paper we survey the evaluation methodologies used in AQG.Based on a sample of 37 papers, our research shows that the systems' development has not been accompanied by similar developments in the methodologies used for the systems' evaluation.Indeed, in the papers we examine here, we find a wide variety of both intrinsic and extrinsic evaluation methodologies.Such diverse evaluation practices make it difficult to reliably compare the quality of different generation systems.Our study suggests that, given the rapidly increasing level of research in the area, a common framework is urgently needed to compare the performance of AQG systems and NLG systems more generally.
Jacopo Amidei, Paul Piwek, Alistair Willis
INLG3
2017 Personalised Query Suggestion for Intranet Search with Temporal User Profiling
abstract
Recent research has shown the usefulness of using collective user interaction data (e.g., query logs) to recommend query modification suggestions for Intranet search. However, most of the query suggestion approaches for Intranet search follow an ``one size fits all'' strategy, whereby different users who submit an identical query would get the same query suggestion list. This is problematic, as even with the same query, different users may have different topics of interest, which may change over time in response to the user's interaction with the system.
Alistair Willis, Udo Kruschwitz, Dawei Song 0001
CHIIR2
2017 Search Personalization with Embeddings
Dat Quoc Nguyen, Mark Johnson 0001, Dawei Song 0001, Alistair Willis
ECIR5
2016 Adverse Drug Reaction Classification With Deep Neural Networks
abstract
We study the problem of detecting sentences describing adverse drug reactions (ADRs) and frame the problem as binary classification. We investigate different neural network (NN) architectures for ADR classification. In particular, we propose two new neural network models, Convolutional Recurrent Neural Network (CRNN) by concatenating convolutional neural networks with recurrent neural networks, and Convolutional Neural Network with Attention (CNNA) by adding attention weights into convolutional neural networks. We evaluate various NN architectures on a Twitter dataset containing informal language and an Adverse Drug Effects (ADE) dataset constructed by sampling from MEDLINE case reports. Experimental results show that all the NN architectures outperform the traditional maximum entropy classifiers trained from n-grams with different weighting strategies considerably on both datasets. On the Twitter dataset, all the NN architectures perform similarly. But on the ADE dataset, CNN performs better than other more complex CNN variants. Nevertheless, CNNA allows the visualisation of attention weights of words when making classification decisions and hence is more appropriate for the extraction of word subsequences describing ADRs.
Trung Huynh, Yulan He 0001, Alistair Willis, Stefan M. Rüger
COLING3
2015 Temporal Latent Topic User Profiles for Search Personalisation
Thanh Tien Vu, Alistair Willis, Son Ngoc Tran, Dawei Song 0001
ECIR2
2014 Improving search personalisation with dynamic group formation
abstract
Recent research has shown that the performance of search engines can be improved by enriching a user's personal profile with information about other users with shared interests. In the existing approaches, groups of similar users are often statically determined, e.g., based on the common documents that users clicked. However, these static grouping methods are query-independent and neglect the fact that users in a group may have different interests with respect to different topics. In this paper, we argue that common interest groups should be dynamically constructed in response to the user's input query. We propose a personalisation framework in which a user profile is enriched using information from other users dynamically grouped with respect to an input query. The experimental results on query logs from a major commercial web search engine demonstrate that our framework improves the performance of the web search engine and also achieves better performance than the static grouping method.
Thanh Tien Vu, Dawei Song 0001, Alistair Willis, Son Ngoc Tran, Jingfei Li
SIGIR3
2012 Speculative requirements: Automatic detection of uncertainty in natural language requirements
abstract
Stakeholders frequently use speculative language when they need to convey their requirements with some degree of uncertainty. Due to the intrinsic vagueness of speculative language, speculative requirements risk being misunderstood, and related uncertainty overlooked, and may benefit from careful treatment in the requirements engineering process. In this paper, we present a linguistically-oriented approach to automatic detection of uncertainty in natural language (NL) requirements. Our approach comprises two stages. First we identify speculative sentences by applying a machine learning algorithm called Conditional Random Fields (CRFs) to identify uncertainty cues. The algorithm exploits a rich set of lexical and syntactic features extracted from requirements sentences. Second, we try to determine the scope of uncertainty. We use a rule-based approach that draws on a set of hand-crafted linguistic heuristics to determine the uncertainty scope with the help of dependency structures present in the sentence parse tree. We report on a series of experiments we conducted to evaluate the performance and usefulness of our system.
Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh
RE4
2011 Analysing anaphoric ambiguity in natural language requirements
Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh
Requir. Eng.4
2010 A Methodology for Automatic Identification of Nocuous Ambiguity
Hui Yang 0004, Anne N. De Roeck, Alistair Willis, Bashar Nuseibeh
COLING3
2010 Using Discovered, Polyphonic Patterns to Filter Computer-generated Music
Tom Collins, Robin C. Laney, Alistair Willis, Paul H. Garthwaite
ICCC3
2010 Automatic detection of nocuous coordination ambiguities in natural language requirements
abstract
Natural language is prevalent in requirements documents. However, ambiguity is an intrinsic phenomenon of natural language, and is therefore present in all such documents. Ambiguity occurs when a sentence can be interpreted differently by different readers. In this paper, we describe an automated approach for characterizing and detecting so-called nocuous ambiguities, which carry a high risk of misunderstanding among different readers. Given a natural language requirements document, sentences that contain specific types of ambiguity are first extracted automatically from the text. A machine learning algorithm is then used to determine whether an ambiguous sentence is nocuous or innocuous, based on a set of heuristics that draw on human judgments, which we collected as training data. We implemented a prototype tool for Nocuous Ambiguity Identification (NAI), in order to illustrate and evaluate our approach. The tool focuses on coordination ambiguity. We report on the results of a set of experiments to assess the performance and usefulness of the approach.
Hui Yang 0004, Alistair Willis, Anne N. De Roeck, Bashar Nuseibeh
ASE2
2010 From XML to XML: The Why and How of Making the Biodiversity Literature Accessible to Researchers
Alistair Willis, David King, David R. Morse, Anton Dil, Chris Lyal, Dave Roberts
LREC1
2010 Extending Nocuous Ambiguity Analysis for Anaphora in Natural Language Requirements
abstract
This paper presents an approach to automatically identify potentially nocuous ambiguities, which occur when text is interpreted differently by different readers of requirements written in natural language. We extract a set of anaphora ambiguities from a range of requirements documents, and collect multiple human judgments on their interpretations. The judgment distribution is used to determine if an ambiguity is nocuous or innocuous. We investigate a number of antecedent preference heuristics that we use to explore aspects of anaphora which may lead a reader to favour a particular interpretation. Using machine learning techniques, we build an automated tool to predict the antecedent preference of noun phrase candidates, which in turn is used to identify nocuous ambiguity. We report on a series of experiments that we conducted to evaluate the performance of our automated system. The results show that the system achieves high recall with a consistent improvement on baseline precision subject to some ambiguity tolerance levels, allowing us to explore and highlight realistic and potentially problematic ambiguities in actual requirements documents.
Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh
RE4
2006 Identifying Nocuous Ambiguities in Natural Language Requirements
abstract
We present a novel technique that automatically alerts authors of requirements to the presence of potentially dangerous ambiguities. We first establish the notion of nocuous ambiguities, which are those that are likely to lead to misunderstandings. We test our approach on coordination ambiguities, which occur when words such as and or are used. Our starting point is a dataset of ambiguous phrases from a requirements corpus and associated human judgements about their interpretation. We then use heuristics, based largely on word distribution information, to automatically replicate these judgements. The heuristics eliminate ambiguities which people interpret easily, leaving the nocuous ones to be analysed and rewritten by hand. We report on a series of experiments that evaluate our heuristics' performance against the human judgements. Many of our heuristics achieve high precision, and recall is greatly increased when they are used in combination
Francis Chantree, Bashar Nuseibeh, Anne N. De Roeck, Alistair Willis
RE4
1999 Two Accounts of Scope Availability and Semantic Underspecification
abstract
We propose a formal system for representing the available readings of sentences displaying quantifier scope ambiguity, in which partial scopes may be expressed.We show that using a theory of scope availability based upon the functionargument structure of a sentence allows a deterministic, polynomial time test for the availability of a reading, while solving the same problem within theories based on the well-formedness of sentences in the meaning language has been shown to be NP-hard.
Alistair Willis, Suresh Manandhar
ACL1