Ronan Cummins

dblp:28/1264 · DBLP profile ↗
← Back
25ranked-venue papers
20as first author
0since 2021 · last 2019
0000-0001-9404-4220ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 21 · 16 first-authorArtificial intelligence and machine learning · 10 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Information retrieval · 100%
Human-computer interaction and pervasive computing
2 papers
Health and well-being technologies · 66% Learning and educational technologies · 34%
Artificial intelligence
1 paper
Learning paradigms · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.732016
A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016
A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015
Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014
Information retrieval › retrieval models
language model
0.522016
A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016
A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015
Information retrieval › evaluation
query performance prediction
0.532014
Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014
Investigating performance predictors using monte carlo simulation and score distribution models · SIGIR 2012
Improved query performance prediction using standard deviation · SIGIR 2011
Information retrieval
evaluation
0.432014
Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014
Investigating performance predictors using monte carlo simulation and score distribution models · SIGIR 2012
Measuring constraint violations in information retrieval · SIGIR 2009
Machine learning › Learning paradigms
multi-task learning
0.212016
Constrained Multi-Task Learning for Automated Essay Scoring · ACL (1) 2016
Information retrieval
query processing
0.212016
A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016
Learning and educational technologies
automated essay scoring
0.212016
Constrained Multi-Task Learning for Automated Essay Scoring · ACL (1) 2016
Information retrieval › retrieval models
term dependency
0.212015
A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015
Information retrieval
retrieval evaluation
0.222013
Which vertical search engines are relevant? · WWW 2013
Evaluating aggregated search pages · SIGIR 2012
Information retrieval › web search
vertical selection
0.222013
Which vertical search engines are relevant? · WWW 2013
Evaluating aggregated search pages · SIGIR 2012
Information retrieval › distributed information retrieval
aggregated search
0.112012
Evaluating aggregated search pages · SIGIR 2012
Health and well-being technologies › mental health technology
digital mental health intervention
0.112019
TIM: A Tool for Gaining Insights into Psychotherapy · WWW 2019
Information retrieval › retrieval models
ad-hoc retrieval
0.112009
Learning in a pairwise term-term proximity framework for information retrieval · SIGIR 2009
Information retrieval › retrieval models
axiomatic analysis
0.112009
Measuring constraint violations in information retrieval · SIGIR 2009
Information retrieval › ranking
ranking model
0.112009
Learning in a pairwise term-term proximity framework for information retrieval · SIGIR 2009
Information retrieval › retrieval models
term weighting
0.112009
Measuring constraint violations in information retrieval · SIGIR 2009
Information retrieval › query reformulation
query expansion
0.112016
A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016
Information retrieval › evaluation
user study
0.012013
Which vertical search engines are relevant? · WWW 2013

Methods — techniques the papers use, named apart from their topics

pairwise preference learning · 0.5multi-task learning · 0.5natural language processing · 0.4machine learning · 0.4document length normalization · 0.2discriminative language modeling · 0.2pólya process · 0.2dirichlet compound multinomial · 0.2parameterised mixture distributions · 0.2average precision inference · 0.2crowdsourcing · 0.2score distribution models · 0.1monte carlo simulation · 0.1evaluation metrics · 0.1
YearPublicationVenuePosition
2019 TIM: A Tool for Gaining Insights into Psychotherapy
abstract
We introduce and demonstrate the usefulness of a tool that automatically annotates therapist utterances in real-time according to the therapeutic role that they perform in an evidence-based psychological dialogue. This is implemented within the context of an on-line service that supports the delivery of one-to-one therapy. When combined with patient outcome measures, this tool allows us to discover the active ingredients in psychotherapy. In particular, we show that particular measures of therapy content are more strongly correlated with patient improvement than others, suggesting that they are a critical part of psychotherapy. As this tool gives us interpretable measures of therapy content, it can enable services to quality control the therapy delivered. Furthermore, we show how specific insights can be presented to the therapist so they can reflect on and improve their practice.
Ronan Cummins, Michael P. Ewbank, Alan Martin, Valentin Tablan, Ana Catarino, Andrew D. Blackwell
WWW1
2016 Constrained Multi-Task Learning for Automated Essay Scoring
abstract
Supervised machine learning models for automated essay scoring (AES) usually require substantial task-specific training data in order to make accurate predictions for a particular writing task.This limitation hinders their utility, and consequently their deployment in real-world settings.In this paper, we overcome this shortcoming using a constrained multi-task pairwisepreference learning approach that enables the data from multiple tasks to be combined effectively.Furthermore, contrary to some recent research, we show that high performance AES systems can be built with little or no task-specific training data.We perform a detailed study of our approach on a publicly available dataset in scenarios where we have varying amounts of task-specific training data and in scenarios where the number of tasks increases.
Ronan Cummins, Meng Zhang 0021, Ted Briscoe
ACL (1)1
2016 A Study of Retrieval Models for Long Documents and Queries in Information Retrieval
abstract
Recent research has shown that long documents are unfairly penalised by a number of current retrieval methods. In this paper, we formally analyse two important but distinct reasons for normalising documents with respect to length, namely verbosity and scope, and discuss the practical implications of not normalising accordingly. We review a number of language modelling approaches and a range of recently developed retrieval methods, and show that most do not correctly model both phenomena, thus limiting their retrieval effectiveness in certain situations. Furthermore, the retrieval characteristics of long natural language queries have not traditionally had the same attention as short keyword queries. We develop a new discriminative query language modelling approach that demonstrates improved performance on long verbose queries by appropriately weighting salient aspects of the query. When combined with query expansion, we show that our new approach yields state-of-the-art performance for long verbose queries.
Ronan Cummins
WWW1
2015 A Pólya Urn Document Language Model for Improved Information Retrieval
abstract
The multinomial language model has been one of the most effective models of retrieval for more than a decade. However, the multinomial distribution does not model one important linguistic phenomenon relating to term dependency—that is, the tendency of a term to repeat itself within a document (i.e., word burstiness). In this article, we model document generation as a random process with reinforcement (a multivariate Pólya process) and develop a Dirichlet compound multinomial language model that captures word burstiness directly. We show that the new reinforced language model can be computed as efficiently as current retrieval models, and with experiments on an extensive set of TREC collections, we show that it significantly outperforms the state-of-the-art language model for a number of standard effectiveness metrics. Experiments also show that the tuning parameter in the proposed model is more robust than that in the multinomial language model. Furthermore, we develop a constraint for the verbosity hypothesis and show that the proposed model adheres to the constraint. Finally, we show that the new language model essentially introduces a measure closely related to idf, which gives theoretical justification for combining the term and document event spaces in tf-idf type schemes.
Ronan Cummins, Jiaul H. Paik, Yuanhua Lv
ACM Trans. Inf. Syst.1
2014 Document Score Distribution Models for Query Performance Inference and Prediction
abstract
Modelling the distribution of document scores returned from an information retrieval (IR) system in response to a query is of both theoretical and practical importance. One of the goals of modelling document scores in this manner is the inference of document relevance. There has been renewed interest of late in modelling document scores using parameterised distributions. Consequently, a number of hypotheses have been proposed to constrain the mixture distribution from which document scores could be drawn. In this article, we show how a standard performance measure (i.e., average precision) can be inferred from a document score distribution using labelled data. We use the accuracy of the inference of average precision as a measure for determining the usefulness of a particular model of document scores. We provide a comprehensive study which shows that certain mixtures of distributions are able to infer average precision more accurately than others. Furthermore, we analyse a number of mixture distributions with regard to the recall-fallout convexity hypothesis and show that the convexity hypothesis is practically useful. Consequently, based on one of the best-performing score-distribution models, we develop some techniques for query-performance prediction (QPP) by automatically estimating the parameters of the document score-distribution model when relevance information is unknown. We present experimental results that outline the benefits of this approach to query-performance prediction.
Ronan Cummins
ACM Trans. Inf. Syst.1
2013 On the reliability and intuitiveness of aggregated search metrics
abstract
Aggregating search results from a variety of diverse verticals such as news, images, videos and Wikipedia into a single interface is a popular web search presentation paradigm. Although several aggregated search (AS) metrics have been proposed to evaluate AS result pages, their properties remain poorly understood. In this paper, we compare the properties of existing AS metrics under the assumptions that (1) queries may have multiple preferred verticals; (2) the likelihood of each vertical preference is available; and (3) the topical relevance assessments of results returned from each vertical is available. We compare a wide range of AS metrics on two test collections. Our main criteria of comparison are (1) discriminative power, which represents the reliability of a metric in comparing the performance of systems, and (2) intuitiveness, which represents how well a metric captures the various key aspects to be measured (i.e. various aspects of a user's perception of AS result pages). Our study shows that the AS metrics that capture key AS components (e.g., vertical selection) have several advantages over other metrics. This work sheds new lights on the further developments and applications of AS metrics.
Ke Zhou 0003, Mounia Lalmas-Roelleke, Tetsuya Sakai, Ronan Cummins, Joemon M. Jose
CIKM4
2013 Which vertical search engines are relevant?
abstract
Aggregating search results from a variety of heterogeneous sources, so-called verticals, such as news, image and video, into a single interface is a popular paradigm in web search. Current approaches that evaluate the effectiveness of aggregated search systems are based on rewarding systems that return highly relevant verticals for a given query, where this relevance is assessed under different assumptions. It is difficult to evaluate or compare those systems without fully understanding the relationship between those underlying assumptions. To address this, we present a formal analysis and a set of extensive user studies to investigate the effects of various assumptions made for assessing query vertical relevance. A total of more than 20,000 assessments on 44 search tasks across 11 verticals are collected through Amazon Mechanical Turk and subsequently analysed. Our results provide insights into various aspects of query vertical relevance and allow us to explain in more depth as well as questioning the evaluation results published in the literature.
Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose
WWW2
2012 On the inference of average precision from score distributions
abstract
Modelling the document scores returned from an IR system for a given query using parameterised score distributions is an area of research that has become more popular in recent years. Score distribution (SD) models are useful for a number of IR tasks. These include data fusion, query performance prediction, determining thresholds in filtering applications, and tasks in the area of distributed retrieval. The inference of performance metrics, such as average precision, from these SD models is an important consideration. In this paper, we study the accuracy of a number of methods of inferring average precision from an SD model.
Ronan Cummins
CIKM1
2012 A constraint to automatically regulate document-length normalisation
abstract
Retrieval functions in information retrieval (IR) are fundamental to the effectiveness of search systems. However, considerable parameter tuning is often needed to increase the effectiveness of the retrieval. Document length normalisation is one such aspect that requires tuning on a per-query and per-collection basis for many retrieval functions.
Ronan Cummins, Colm O'Riordan
CIKM1
2012 Evaluating reward and risk for vertical selection
abstract
The aggregation of search results from heterogeneous verticals (news, videos, blogs, etc) has become an important consideration in search. When aiming to select suitable verticals, from which items are selected to be shown along with the standard "ten blue links", there exists the potential to both help (selecting relevant verticals) and harm (selecting irrelevant verticals) the existing result set.
Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose
CIKM2
2012 On Theoretically Valid Score Distributions in Information Retrieval
Ronan Cummins, Colm O'Riordan
ECIR1
2012 Assessing and Predicting Vertical Intent for Web Queries
Ke Zhou 0003, Ronan Cummins, Martin Halvey, Mounia Lalmas-Roelleke, Joemon M. Jose
ECIR2
2012 Investigating performance predictors using monte carlo simulation and score distribution models
abstract
The standard deviation of scores in the top k documents of a ranked list has been shown to be significantly correlated with average precision and has been the basis of a number of query performance predictors. In this paper, we outline two hypotheses that aid in understanding this correlation. Using score distribution (SD) models with known parameters, we create a large number of document rankings using Monte Carlo simulation to test the validity of these hypotheses.
Ronan Cummins
SIGIR1
2012 Evaluating aggregated search pages
abstract
Aggregating search results from a variety of heterogeneous sources or verticals such as news, image and video into a single interface is a popular paradigm in web search. Although various approaches exist for selecting relevant verticals or optimising the aggregated search result page, evaluating the quality of an aggregated page is an open question.
Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose
SIGIR2
2011 The Limits of Retrieval Effectiveness
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan
ECIR1
2011 Improved query performance prediction using standard deviation
abstract
Query performance prediction (QPP) is an important task in information retrieval (IR). In this paper, we (1) develop a new predictor based on the standard deviation of scores in a variable length ranked list, and (2) we show that this new predictor outperforms state-of-the-art approaches without the need for tuning.
Ronan Cummins, Joemon M. Jose, Colm O'Riordan
SIGIR1
2011 Navigating the User Query Space
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan, Joemon M. Jose
SPIRE1
2010 Examining the information retrieval process from an inductive perspective
abstract
Term-weighting functions derived from various models of retrieval aim to model human notions of relevance more accurately. However, there is a lack of analysis of the sources of evidence from which important features of these term weighting schemes originate. In general, features pertaining to these term-weighting schemes can be collected from (1) the document, (2) the entire collection and (3) the query. In this work, we perform an empirical analysis to determine the increase in effectiveness as information from these three different sources becomes more accurate.
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan
CIKM1
2010 Learning Aggregation Functions for Expert Search
abstract
Machine learning techniques are increasingly being applied to problems in the domain of information retrieval and text mining. In this paper we present an application of evolutionary computation to the area of expert search. Expert search in the context of enterprise information systems deals with the problem of finding and ranking candidate experts given an information need (query). A difficult problem in the area of expert search is finding relevant information given an information need and associating that information with a potential expert.
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan
ECAI1
2009 Learning in a pairwise term-term proximity framework for information retrieval
abstract
Traditional ad hoc retrieval models do not take into account the closeness or proximity of terms. Document scores in these models are primarily based on the occurrences or nonoccurrences of query-terms considered independently of each other. Intuitively, documents in which query-terms occur closer together should be ranked higher than documents in which the query-terms appear far apart. \nThis paper outlines several term-term proximity measures and develops an intuitive framework in which they can be used to fully model the proximity of all query-terms for a particular topic. As useful proximity functions may be constructed from many proximity measures, we use a learning approach to combine proximity measures to develop a useful proximity function in the framework. An evaluation of the best proximity functions show that there is a significant improvement over the baseline ad hoc retrieval model and over other more recent methods that employ the use of single proximity measures.
Ronan Cummins, Colm O'Riordan
SIGIR1
2009 Measuring constraint violations in information retrieval
abstract
Recently, an inductive approach to modelling term-weighting function correctness has provided a number of axioms (constraints), to which all good term-weighting functions should adhere. These constraints have been shown to be theoretically and empirically sound in a number of works. It has been shown that when a term-weighting function breaks one or more of the constraints, it typically indicates sub-optimality of that function. This elegant inductive approach may more accurately model the human process of determining the relevance a document. It is intuitive that a person's notion of relevance changes as terms that are either on or off-topic are encountered in a given document. Ultimately, it would be desirable to be able to mathematically determine the performance of term-weighting functions without the need for test collections.
Ronan Cummins, Colm O'Riordan
SIGIR1
2007 Using genetic programming for information retrieval: local and global query expansion
abstract
This poster presents results for two approaches using Genetic Programming (GP) to overcome the problem of term mismatch in Information Retrieval (IR). We use automatic query expansion techniques which add terms to a user's initial query in the hope that these words better describe the information need and ultimately return more relevant documents to the user.
Ronan Cummins, Colm O'Riordan
GECCO1
2006 Weighting in Information Retrieval Using Genetic Programming: A Three Stage Process
Ronan Cummins, Colm O'Riordan
ECAI1
2006 Evolving local and global weighting schemes in information retrieval
Ronan Cummins, Colm O'Riordan
Inf. Retr.1
2005 An evaluation of evolved term-weighting schemes in information retrieval
abstract
This paper presents an evaluation of evolved term-weighting schemes on short, medium and long TREC queries. A previously evolved global (collection-wide) term-weighting scheme is evaluated on unseen TREC data and is shown to increase mean average precision over idf. A local (within-document) evolved term-weighting scheme is presented which is dependent on the best performing global scheme. The full evolved scheme (i.e. the combined local and global scheme) is compared to both the BM25 scheme and the Pivoted Normalisation scheme.Our results show that the local evolved solution does not perform well on some collections due to its document normalisation properties and we conclude that Okapi-tf can be tuned to interact effectively with the evolved global weighting scheme presented and increase mean average precision over the standard BM25 scheme.
Ronan Cummins, Colm O'Riordan
CIKM1