VLDB 2026 Research / reviewers in the wild / expert
Alistair Willis
dblp:67/4642
· DBLP profile ↗
20ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0001-7233-8623ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 5Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Recommender systems · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 91% Empirical software engineering · 9% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › personalized search
group-based personalization |
0.2 | 1 | 2014 | Improving search personalisation with dynamic group formation · SIGIR 2014 |
Information retrieval
personalized search |
0.2 | 1 | 2014 | Improving search personalisation with dynamic group formation · SIGIR 2014 |
Recommender systems
user modeling |
0.2 | 1 | 2014 | Improving search personalisation with dynamic group formation · SIGIR 2014 |
Requirements engineering and software design › requirements quality
ambiguity detection |
0.1 | 1 | 2010 | Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010 |
Requirements engineering and software design › requirements specification
natural language requirements |
0.1 | 1 | 2010 | Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010 |
Requirements engineering and software design
requirements quality |
0.1 | 1 | 2010 | Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010 |
Empirical software engineering
developer studies |
0.0 | 1 | 2010 | Automatic detection of nocuous coordination ambiguities in natural language requirements · ASE 2010 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic in computer science › logical foundations
formal semantics |
0.0 | 1 | 1999 | Two Accounts of Scope Availability and Semantic Underspecification · ACL 1999 |
Algorithms and data structures
polynomial-time algorithms |
0.0 | 1 | 1999 | Two Accounts of Scope Availability and Semantic Underspecification · ACL 1999 |
Methods — techniques the papers use, named apart from their topics
machine learning · 0.1heuristics · 0.1underspecification · 0.0formal semantics · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Identifying Annotator Bias: A new IRT-based method for bias identificationabstractA basic step in any annotation effort is the measurement of the Inter Annotator Agreement (IAA).An important factor that can affect the IAA is the presence of annotator bias.In this paper we introduce a new interpretation and application of the Item Response Theory (IRT) to detect annotators' bias.Our interpretation of IRT offers an original bias identification method that can be used to compare annotators' bias and characterise annotation disagreement.Our method can be used to spot outlier annotators, improve annotation guidelines and provide a better picture of the annotation reliability.Additionally, because scales for IAA interpretation are not generally agreed upon, our bias identification method is valuable as a complement to the IAA value which can help with understanding the annotation disagreement. Jacopo Amidei, Paul Piwek, Alistair Willis |
COLING | 3 |
| 2020 | Developing Students' Written Communication Skills with Jupyter NotebooksabstractWritten communication skills are considered to be highly desirable in computing graduates. However, many computing students do not have a background in which these skills have been developed, and the skills are often not well addressed within a computing curriculum. For some multidisciplinary areas, such as data science, the range of potential stakeholders makes the need for communications skills all the greater. As interest in data science increases and the technical skills of the area are in ever higher demand, understanding effective teaching and learning of these interdisciplinary aspects is receiving significant attention by academics, industry and government in an effort to address the digital skills gap. In this paper, we report on the experience of adapting a final year data science module in an undergraduate computing curriculum to help develop the skills needed for writing extended reports. From its inception, the module has used Jupyter notebooks to develop the students' skills in the coding aspects of the module. However, over several presentations, we have investigated how the cell-based structure of the notebooks can be exploited to improve the students' understanding of how to structure a report on a data investigation. We have increasingly designed the assessment for the module to take advantage of the learning affordances of Jupyter notebooks to support both raw data analysis and effective report writing. We reflect on the lessons learned from these changes to the assessment model, and the students' responses to the changes. Alistair Willis, Patricia Charlton, Tony Hirst |
SIGCSE | 1 |
| 2019 | Agreement is overrated: A plea for correlation to assess human evaluation reliabilityabstractInter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability.According to existing scales of IAA interpretationsee, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) -most data collected for NLG evaluation fail the reliability test.We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study).Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. ( 2018b), such phenomena can be explained in terms of irreducible human language variability.Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability.Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability.This is illustrated using the three case studies. Jacopo Amidei, Paul Piwek, Alistair Willis |
INLG | 3 |
| 2019 | The use of rating and Likert scales in Natural Language Generation human evaluation tasks: A review and some recommendationsabstractRating and Likert scales are widely used in evaluation experiments to measure the quality of Natural Language Generation (NLG) systems.We review the use of rating and Likert scales for NLG evaluation tasks published in NLG specialized conferences over the last ten years (135 papers in total).Our analysis brings to light a number of deviations from good practice in their use.We conclude with some recommendations about the use of such scales.Our aim is to encourage the appropriate use of evaluation methodologies in the NLG community. Jacopo Amidei, Paul Piwek, Alistair Willis |
INLG | 3 |
| 2018 | Rethinking the Agreement in Human Evaluation TasksabstractHuman evaluations are broadly thought to be more valuable the higher the inter-annotator agreement. In this paper we examine this idea. We will describe our experiments and analysis within the area of Automatic Question Generation. Our experiments show how annotators diverge in language annotation tasks due to a range of ineliminable factors. For this reason, we believe that annotation schemes for natural language generation tasks that are aimed at evaluating language quality need to be treated with great care. In particular, an unchecked focus on reduction of disagreement among annotators runs the danger of creating generation goals that reward output that is more distant from, rather than closer to, natural human-like language. We conclude the paper by suggesting a new approach to the use of the agreement metrics in natural language generation evaluation tasks. Jacopo Amidei, Paul Piwek, Alistair Willis |
COLING | 3 |
| 2018 | Evaluation methodologies in Automatic Question Generation 2013-2018abstractIn the last few years Automatic Question Generation (AQG) has attracted increasing interest.In this paper we survey the evaluation methodologies used in AQG.Based on a sample of 37 papers, our research shows that the systems' development has not been accompanied by similar developments in the methodologies used for the systems' evaluation.Indeed, in the papers we examine here, we find a wide variety of both intrinsic and extrinsic evaluation methodologies.Such diverse evaluation practices make it difficult to reliably compare the quality of different generation systems.Our study suggests that, given the rapidly increasing level of research in the area, a common framework is urgently needed to compare the performance of AQG systems and NLG systems more generally. Jacopo Amidei, Paul Piwek, Alistair Willis |
INLG | 3 |
| 2017 | Personalised Query Suggestion for Intranet Search with Temporal User ProfilingabstractRecent research has shown the usefulness of using collective user interaction data (e.g., query logs) to recommend query modification suggestions for Intranet search. However, most of the query suggestion approaches for Intranet search follow an ``one size fits all'' strategy, whereby different users who submit an identical query would get the same query suggestion list. This is problematic, as even with the same query, different users may have different topics of interest, which may change over time in response to the user's interaction with the system. Alistair Willis, Udo Kruschwitz, Dawei Song 0001 |
CHIIR | 2 |
| 2017 | Search Personalization with Embeddings
Dat Quoc Nguyen, Mark Johnson 0001, Dawei Song 0001, Alistair Willis |
ECIR | 5 |
| 2016 | Adverse Drug Reaction Classification With Deep Neural NetworksabstractWe study the problem of detecting sentences describing adverse drug reactions (ADRs) and frame the problem as binary classification. We investigate different neural network (NN) architectures for ADR classification. In particular, we propose two new neural network models, Convolutional Recurrent Neural Network (CRNN) by concatenating convolutional neural networks with recurrent neural networks, and Convolutional Neural Network with Attention (CNNA) by adding attention weights into convolutional neural networks. We evaluate various NN architectures on a Twitter dataset containing informal language and an Adverse Drug Effects (ADE) dataset constructed by sampling from MEDLINE case reports. Experimental results show that all the NN architectures outperform the traditional maximum entropy classifiers trained from n-grams with different weighting strategies considerably on both datasets. On the Twitter dataset, all the NN architectures perform similarly. But on the ADE dataset, CNN performs better than other more complex CNN variants. Nevertheless, CNNA allows the visualisation of attention weights of words when making classification decisions and hence is more appropriate for the extraction of word subsequences describing ADRs. Trung Huynh, Yulan He 0001, Alistair Willis, Stefan M. Rüger |
COLING | 3 |
| 2015 | Temporal Latent Topic User Profiles for Search Personalisation
Thanh Tien Vu, Alistair Willis, Son Ngoc Tran, Dawei Song 0001 |
ECIR | 2 |
| 2014 | Improving search personalisation with dynamic group formationabstractRecent research has shown that the performance of search engines can be improved by enriching a user's personal profile with information about other users with shared interests. In the existing approaches, groups of similar users are often statically determined, e.g., based on the common documents that users clicked. However, these static grouping methods are query-independent and neglect the fact that users in a group may have different interests with respect to different topics. In this paper, we argue that common interest groups should be dynamically constructed in response to the user's input query. We propose a personalisation framework in which a user profile is enriched using information from other users dynamically grouped with respect to an input query. The experimental results on query logs from a major commercial web search engine demonstrate that our framework improves the performance of the web search engine and also achieves better performance than the static grouping method. Thanh Tien Vu, Dawei Song 0001, Alistair Willis, Son Ngoc Tran, Jingfei Li |
SIGIR | 3 |
| 2012 | Speculative requirements: Automatic detection of uncertainty in natural language requirementsabstractStakeholders frequently use speculative language when they need to convey their requirements with some degree of uncertainty. Due to the intrinsic vagueness of speculative language, speculative requirements risk being misunderstood, and related uncertainty overlooked, and may benefit from careful treatment in the requirements engineering process. In this paper, we present a linguistically-oriented approach to automatic detection of uncertainty in natural language (NL) requirements. Our approach comprises two stages. First we identify speculative sentences by applying a machine learning algorithm called Conditional Random Fields (CRFs) to identify uncertainty cues. The algorithm exploits a rich set of lexical and syntactic features extracted from requirements sentences. Second, we try to determine the scope of uncertainty. We use a rule-based approach that draws on a set of hand-crafted linguistic heuristics to determine the uncertainty scope with the help of dependency structures present in the sentence parse tree. We report on a series of experiments we conducted to evaluate the performance and usefulness of our system. Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh |
RE | 4 |
| 2011 | Analysing anaphoric ambiguity in natural language requirements
Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh |
Requir. Eng. | 4 |
| 2010 | A Methodology for Automatic Identification of Nocuous Ambiguity
Hui Yang 0004, Anne N. De Roeck, Alistair Willis, Bashar Nuseibeh |
COLING | 3 |
| 2010 | Using Discovered, Polyphonic Patterns to Filter Computer-generated Music
Tom Collins, Robin C. Laney, Alistair Willis, Paul H. Garthwaite |
ICCC | 3 |
| 2010 | Automatic detection of nocuous coordination ambiguities in natural language requirementsabstractNatural language is prevalent in requirements documents. However, ambiguity is an intrinsic phenomenon of natural language, and is therefore present in all such documents. Ambiguity occurs when a sentence can be interpreted differently by different readers. In this paper, we describe an automated approach for characterizing and detecting so-called nocuous ambiguities, which carry a high risk of misunderstanding among different readers. Given a natural language requirements document, sentences that contain specific types of ambiguity are first extracted automatically from the text. A machine learning algorithm is then used to determine whether an ambiguous sentence is nocuous or innocuous, based on a set of heuristics that draw on human judgments, which we collected as training data. We implemented a prototype tool for Nocuous Ambiguity Identification (NAI), in order to illustrate and evaluate our approach. The tool focuses on coordination ambiguity. We report on the results of a set of experiments to assess the performance and usefulness of the approach. Hui Yang 0004, Alistair Willis, Anne N. De Roeck, Bashar Nuseibeh |
ASE | 2 |
| 2010 | From XML to XML: The Why and How of Making the Biodiversity Literature Accessible to Researchers
Alistair Willis, David King, David R. Morse, Anton Dil, Chris Lyal, Dave Roberts |
LREC | 1 |
| 2010 | Extending Nocuous Ambiguity Analysis for Anaphora in Natural Language RequirementsabstractThis paper presents an approach to automatically identify potentially nocuous ambiguities, which occur when text is interpreted differently by different readers of requirements written in natural language. We extract a set of anaphora ambiguities from a range of requirements documents, and collect multiple human judgments on their interpretations. The judgment distribution is used to determine if an ambiguity is nocuous or innocuous. We investigate a number of antecedent preference heuristics that we use to explore aspects of anaphora which may lead a reader to favour a particular interpretation. Using machine learning techniques, we build an automated tool to predict the antecedent preference of noun phrase candidates, which in turn is used to identify nocuous ambiguity. We report on a series of experiments that we conducted to evaluate the performance of our automated system. The results show that the system achieves high recall with a consistent improvement on baseline precision subject to some ambiguity tolerance levels, allowing us to explore and highlight realistic and potentially problematic ambiguities in actual requirements documents. Hui Yang 0004, Anne N. De Roeck, Vincenzo Gervasi, Alistair Willis, Bashar Nuseibeh |
RE | 4 |
| 2006 | Identifying Nocuous Ambiguities in Natural Language RequirementsabstractWe present a novel technique that automatically alerts authors of requirements to the presence of potentially dangerous ambiguities. We first establish the notion of nocuous ambiguities, which are those that are likely to lead to misunderstandings. We test our approach on coordination ambiguities, which occur when words such as and or are used. Our starting point is a dataset of ambiguous phrases from a requirements corpus and associated human judgements about their interpretation. We then use heuristics, based largely on word distribution information, to automatically replicate these judgements. The heuristics eliminate ambiguities which people interpret easily, leaving the nocuous ones to be analysed and rewritten by hand. We report on a series of experiments that evaluate our heuristics' performance against the human judgements. Many of our heuristics achieve high precision, and recall is greatly increased when they are used in combination Francis Chantree, Bashar Nuseibeh, Anne N. De Roeck, Alistair Willis |
RE | 4 |
| 1999 | Two Accounts of Scope Availability and Semantic UnderspecificationabstractWe propose a formal system for representing the available readings of sentences displaying quantifier scope ambiguity, in which partial scopes may be expressed.We show that using a theory of scope availability based upon the functionargument structure of a sentence allows a deterministic, polynomial time test for the availability of a reading, while solving the same problem within theories based on the well-formedness of sentences in the meaning language has been shown to be NP-hard. Alistair Willis, Suresh Manandhar |
ACL | 1 |