VLDB 2026 Research / reviewers in the wild / expert
Ronan Cummins
dblp:28/1264
· DBLP profile ↗
25ranked-venue papers
20as first author
0since 2021 · last 2019
0000-0001-9404-4220ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 21 · 16 first-authorArtificial intelligence and machine learning · 10 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Information retrieval · 100% | |
| Human-computer interaction and pervasive computing
2 papers |
Health and well-being technologies · 66% Learning and educational technologies · 34% | |
| Artificial intelligence
1 paper |
Learning paradigms · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
0.7 | 3 | 2016 | A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016 A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015 Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014 |
Information retrieval › retrieval models
language model |
0.5 | 2 | 2016 | A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016 A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015 |
Information retrieval › evaluation
query performance prediction |
0.5 | 3 | 2014 | Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014 Investigating performance predictors using monte carlo simulation and score distribution models · SIGIR 2012 Improved query performance prediction using standard deviation · SIGIR 2011 |
Information retrieval
evaluation |
0.4 | 3 | 2014 | Document Score Distribution Models for Query Performance Inference and Prediction · ACM Trans. Inf. Syst. 2014 Investigating performance predictors using monte carlo simulation and score distribution models · SIGIR 2012 Measuring constraint violations in information retrieval · SIGIR 2009 |
Machine learning › Learning paradigms
multi-task learning |
0.2 | 1 | 2016 | Constrained Multi-Task Learning for Automated Essay Scoring · ACL (1) 2016 |
Information retrieval
query processing |
0.2 | 1 | 2016 | A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016 |
Learning and educational technologies
automated essay scoring |
0.2 | 1 | 2016 | Constrained Multi-Task Learning for Automated Essay Scoring · ACL (1) 2016 |
Information retrieval › retrieval models
term dependency |
0.2 | 1 | 2015 | A Pólya Urn Document Language Model for Improved Information Retrieval · ACM Trans. Inf. Syst. 2015 |
Information retrieval
retrieval evaluation |
0.2 | 2 | 2013 | Which vertical search engines are relevant? · WWW 2013 Evaluating aggregated search pages · SIGIR 2012 |
Information retrieval › web search
vertical selection |
0.2 | 2 | 2013 | Which vertical search engines are relevant? · WWW 2013 Evaluating aggregated search pages · SIGIR 2012 |
Information retrieval › distributed information retrieval
aggregated search |
0.1 | 1 | 2012 | Evaluating aggregated search pages · SIGIR 2012 |
Health and well-being technologies › mental health technology
digital mental health intervention |
0.1 | 1 | 2019 | TIM: A Tool for Gaining Insights into Psychotherapy · WWW 2019 |
Information retrieval › retrieval models
ad-hoc retrieval |
0.1 | 1 | 2009 | Learning in a pairwise term-term proximity framework for information retrieval · SIGIR 2009 |
Information retrieval › retrieval models
axiomatic analysis |
0.1 | 1 | 2009 | Measuring constraint violations in information retrieval · SIGIR 2009 |
Information retrieval › ranking
ranking model |
0.1 | 1 | 2009 | Learning in a pairwise term-term proximity framework for information retrieval · SIGIR 2009 |
Information retrieval › retrieval models
term weighting |
0.1 | 1 | 2009 | Measuring constraint violations in information retrieval · SIGIR 2009 |
Information retrieval › query reformulation
query expansion |
0.1 | 1 | 2016 | A Study of Retrieval Models for Long Documents and Queries in Information Retrieval · WWW 2016 |
Information retrieval › evaluation
user study |
0.0 | 1 | 2013 | Which vertical search engines are relevant? · WWW 2013 |
Methods — techniques the papers use, named apart from their topics
pairwise preference learning · 0.5multi-task learning · 0.5natural language processing · 0.4machine learning · 0.4document length normalization · 0.2discriminative language modeling · 0.2pólya process · 0.2dirichlet compound multinomial · 0.2parameterised mixture distributions · 0.2average precision inference · 0.2crowdsourcing · 0.2score distribution models · 0.1monte carlo simulation · 0.1evaluation metrics · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | TIM: A Tool for Gaining Insights into PsychotherapyabstractWe introduce and demonstrate the usefulness of a tool that automatically annotates therapist utterances in real-time according to the therapeutic role that they perform in an evidence-based psychological dialogue. This is implemented within the context of an on-line service that supports the delivery of one-to-one therapy. When combined with patient outcome measures, this tool allows us to discover the active ingredients in psychotherapy. In particular, we show that particular measures of therapy content are more strongly correlated with patient improvement than others, suggesting that they are a critical part of psychotherapy. As this tool gives us interpretable measures of therapy content, it can enable services to quality control the therapy delivered. Furthermore, we show how specific insights can be presented to the therapist so they can reflect on and improve their practice. Ronan Cummins, Michael P. Ewbank, Alan Martin, Valentin Tablan, Ana Catarino, Andrew D. Blackwell |
WWW | 1 |
| 2016 | Constrained Multi-Task Learning for Automated Essay ScoringabstractSupervised machine learning models for automated essay scoring (AES) usually require substantial task-specific training data in order to make accurate predictions for a particular writing task.This limitation hinders their utility, and consequently their deployment in real-world settings.In this paper, we overcome this shortcoming using a constrained multi-task pairwisepreference learning approach that enables the data from multiple tasks to be combined effectively.Furthermore, contrary to some recent research, we show that high performance AES systems can be built with little or no task-specific training data.We perform a detailed study of our approach on a publicly available dataset in scenarios where we have varying amounts of task-specific training data and in scenarios where the number of tasks increases. Ronan Cummins, Meng Zhang 0021, Ted Briscoe |
ACL (1) | 1 |
| 2016 | A Study of Retrieval Models for Long Documents and Queries in Information RetrievalabstractRecent research has shown that long documents are unfairly penalised by a number of current retrieval methods. In this paper, we formally analyse two important but distinct reasons for normalising documents with respect to length, namely verbosity and scope, and discuss the practical implications of not normalising accordingly. We review a number of language modelling approaches and a range of recently developed retrieval methods, and show that most do not correctly model both phenomena, thus limiting their retrieval effectiveness in certain situations. Furthermore, the retrieval characteristics of long natural language queries have not traditionally had the same attention as short keyword queries. We develop a new discriminative query language modelling approach that demonstrates improved performance on long verbose queries by appropriately weighting salient aspects of the query. When combined with query expansion, we show that our new approach yields state-of-the-art performance for long verbose queries. Ronan Cummins |
WWW | 1 |
| 2015 | A Pólya Urn Document Language Model for Improved Information RetrievalabstractThe multinomial language model has been one of the most effective models of retrieval for more than a decade. However, the multinomial distribution does not model one important linguistic phenomenon relating to term dependency—that is, the tendency of a term to repeat itself within a document (i.e., word burstiness). In this article, we model document generation as a random process with reinforcement (a multivariate Pólya process) and develop a Dirichlet compound multinomial language model that captures word burstiness directly. We show that the new reinforced language model can be computed as efficiently as current retrieval models, and with experiments on an extensive set of TREC collections, we show that it significantly outperforms the state-of-the-art language model for a number of standard effectiveness metrics. Experiments also show that the tuning parameter in the proposed model is more robust than that in the multinomial language model. Furthermore, we develop a constraint for the verbosity hypothesis and show that the proposed model adheres to the constraint. Finally, we show that the new language model essentially introduces a measure closely related to idf, which gives theoretical justification for combining the term and document event spaces in tf-idf type schemes. Ronan Cummins, Jiaul H. Paik, Yuanhua Lv |
ACM Trans. Inf. Syst. | 1 |
| 2014 | Document Score Distribution Models for Query Performance Inference and PredictionabstractModelling the distribution of document scores returned from an information retrieval (IR) system in response to a query is of both theoretical and practical importance. One of the goals of modelling document scores in this manner is the inference of document relevance. There has been renewed interest of late in modelling document scores using parameterised distributions. Consequently, a number of hypotheses have been proposed to constrain the mixture distribution from which document scores could be drawn. In this article, we show how a standard performance measure (i.e., average precision) can be inferred from a document score distribution using labelled data. We use the accuracy of the inference of average precision as a measure for determining the usefulness of a particular model of document scores. We provide a comprehensive study which shows that certain mixtures of distributions are able to infer average precision more accurately than others. Furthermore, we analyse a number of mixture distributions with regard to the recall-fallout convexity hypothesis and show that the convexity hypothesis is practically useful. Consequently, based on one of the best-performing score-distribution models, we develop some techniques for query-performance prediction (QPP) by automatically estimating the parameters of the document score-distribution model when relevance information is unknown. We present experimental results that outline the benefits of this approach to query-performance prediction. Ronan Cummins |
ACM Trans. Inf. Syst. | 1 |
| 2013 | On the reliability and intuitiveness of aggregated search metricsabstractAggregating search results from a variety of diverse verticals such as news, images, videos and Wikipedia into a single interface is a popular web search presentation paradigm. Although several aggregated search (AS) metrics have been proposed to evaluate AS result pages, their properties remain poorly understood. In this paper, we compare the properties of existing AS metrics under the assumptions that (1) queries may have multiple preferred verticals; (2) the likelihood of each vertical preference is available; and (3) the topical relevance assessments of results returned from each vertical is available. We compare a wide range of AS metrics on two test collections. Our main criteria of comparison are (1) discriminative power, which represents the reliability of a metric in comparing the performance of systems, and (2) intuitiveness, which represents how well a metric captures the various key aspects to be measured (i.e. various aspects of a user's perception of AS result pages). Our study shows that the AS metrics that capture key AS components (e.g., vertical selection) have several advantages over other metrics. This work sheds new lights on the further developments and applications of AS metrics. Ke Zhou 0003, Mounia Lalmas-Roelleke, Tetsuya Sakai, Ronan Cummins, Joemon M. Jose |
CIKM | 4 |
| 2013 | Which vertical search engines are relevant?abstractAggregating search results from a variety of heterogeneous sources, so-called verticals, such as news, image and video, into a single interface is a popular paradigm in web search. Current approaches that evaluate the effectiveness of aggregated search systems are based on rewarding systems that return highly relevant verticals for a given query, where this relevance is assessed under different assumptions. It is difficult to evaluate or compare those systems without fully understanding the relationship between those underlying assumptions. To address this, we present a formal analysis and a set of extensive user studies to investigate the effects of various assumptions made for assessing query vertical relevance. A total of more than 20,000 assessments on 44 search tasks across 11 verticals are collected through Amazon Mechanical Turk and subsequently analysed. Our results provide insights into various aspects of query vertical relevance and allow us to explain in more depth as well as questioning the evaluation results published in the literature. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
WWW | 2 |
| 2012 | On the inference of average precision from score distributionsabstractModelling the document scores returned from an IR system for a given query using parameterised score distributions is an area of research that has become more popular in recent years. Score distribution (SD) models are useful for a number of IR tasks. These include data fusion, query performance prediction, determining thresholds in filtering applications, and tasks in the area of distributed retrieval. The inference of performance metrics, such as average precision, from these SD models is an important consideration. In this paper, we study the accuracy of a number of methods of inferring average precision from an SD model. Ronan Cummins |
CIKM | 1 |
| 2012 | A constraint to automatically regulate document-length normalisationabstractRetrieval functions in information retrieval (IR) are fundamental to the effectiveness of search systems. However, considerable parameter tuning is often needed to increase the effectiveness of the retrieval. Document length normalisation is one such aspect that requires tuning on a per-query and per-collection basis for many retrieval functions. Ronan Cummins, Colm O'Riordan |
CIKM | 1 |
| 2012 | Evaluating reward and risk for vertical selectionabstractThe aggregation of search results from heterogeneous verticals (news, videos, blogs, etc) has become an important consideration in search. When aiming to select suitable verticals, from which items are selected to be shown along with the standard "ten blue links", there exists the potential to both help (selecting relevant verticals) and harm (selecting irrelevant verticals) the existing result set. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
CIKM | 2 |
| 2012 | On Theoretically Valid Score Distributions in Information Retrieval
Ronan Cummins, Colm O'Riordan |
ECIR | 1 |
| 2012 | Assessing and Predicting Vertical Intent for Web Queries
Ke Zhou 0003, Ronan Cummins, Martin Halvey, Mounia Lalmas-Roelleke, Joemon M. Jose |
ECIR | 2 |
| 2012 | Investigating performance predictors using monte carlo simulation and score distribution modelsabstractThe standard deviation of scores in the top k documents of a ranked list has been shown to be significantly correlated with average precision and has been the basis of a number of query performance predictors. In this paper, we outline two hypotheses that aid in understanding this correlation. Using score distribution (SD) models with known parameters, we create a large number of document rankings using Monte Carlo simulation to test the validity of these hypotheses. Ronan Cummins |
SIGIR | 1 |
| 2012 | Evaluating aggregated search pagesabstractAggregating search results from a variety of heterogeneous sources or verticals such as news, image and video into a single interface is a popular paradigm in web search. Although various approaches exist for selecting relevant verticals or optimising the aggregated search result page, evaluating the quality of an aggregated page is an open question. Ke Zhou 0003, Ronan Cummins, Mounia Lalmas-Roelleke, Joemon M. Jose |
SIGIR | 2 |
| 2011 | The Limits of Retrieval Effectiveness
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan |
ECIR | 1 |
| 2011 | Improved query performance prediction using standard deviationabstractQuery performance prediction (QPP) is an important task in information retrieval (IR). In this paper, we (1) develop a new predictor based on the standard deviation of scores in a variable length ranked list, and (2) we show that this new predictor outperforms state-of-the-art approaches without the need for tuning. Ronan Cummins, Joemon M. Jose, Colm O'Riordan |
SIGIR | 1 |
| 2011 | Navigating the User Query Space
Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan, Joemon M. Jose |
SPIRE | 1 |
| 2010 | Examining the information retrieval process from an inductive perspectiveabstractTerm-weighting functions derived from various models of retrieval aim to model human notions of relevance more accurately. However, there is a lack of analysis of the sources of evidence from which important features of these term weighting schemes originate. In general, features pertaining to these term-weighting schemes can be collected from (1) the document, (2) the entire collection and (3) the query. In this work, we perform an empirical analysis to determine the increase in effectiveness as information from these three different sources becomes more accurate. Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan |
CIKM | 1 |
| 2010 | Learning Aggregation Functions for Expert SearchabstractMachine learning techniques are increasingly being applied to problems in the domain of information retrieval and text mining. In this paper we present an application of evolutionary computation to the area of expert search. Expert search in the context of enterprise information systems deals with the problem of finding and ranking candidate experts given an information need (query). A difficult problem in the area of expert search is finding relevant information given an information need and associating that information with a potential expert. Ronan Cummins, Mounia Lalmas-Roelleke, Colm O'Riordan |
ECAI | 1 |
| 2009 | Learning in a pairwise term-term proximity framework for information retrievalabstractTraditional ad hoc retrieval models do not take into account the closeness or proximity of terms. Document scores in these models are primarily based on the occurrences or nonoccurrences of query-terms considered independently of each other. Intuitively, documents in which query-terms occur closer together should be ranked higher than documents in which the query-terms appear far apart. \nThis paper outlines several term-term proximity measures and develops an intuitive framework in which they can be used to fully model the proximity of all query-terms for a particular topic. As useful proximity functions may be constructed from many proximity measures, we use a learning approach to combine proximity measures to develop a useful proximity function in the framework. An evaluation of the best proximity functions show that there is a significant improvement over the baseline ad hoc retrieval model and over other more recent methods that employ the use of single proximity measures. Ronan Cummins, Colm O'Riordan |
SIGIR | 1 |
| 2009 | Measuring constraint violations in information retrievalabstractRecently, an inductive approach to modelling term-weighting function correctness has provided a number of axioms (constraints), to which all good term-weighting functions should adhere. These constraints have been shown to be theoretically and empirically sound in a number of works. It has been shown that when a term-weighting function breaks one or more of the constraints, it typically indicates sub-optimality of that function. This elegant inductive approach may more accurately model the human process of determining the relevance a document. It is intuitive that a person's notion of relevance changes as terms that are either on or off-topic are encountered in a given document. Ultimately, it would be desirable to be able to mathematically determine the performance of term-weighting functions without the need for test collections. Ronan Cummins, Colm O'Riordan |
SIGIR | 1 |
| 2007 | Using genetic programming for information retrieval: local and global query expansionabstractThis poster presents results for two approaches using Genetic Programming (GP) to overcome the problem of term mismatch in Information Retrieval (IR). We use automatic query expansion techniques which add terms to a user's initial query in the hope that these words better describe the information need and ultimately return more relevant documents to the user. Ronan Cummins, Colm O'Riordan |
GECCO | 1 |
| 2006 | Weighting in Information Retrieval Using Genetic Programming: A Three Stage Process
Ronan Cummins, Colm O'Riordan |
ECAI | 1 |
| 2006 | Evolving local and global weighting schemes in information retrieval
Ronan Cummins, Colm O'Riordan |
Inf. Retr. | 1 |
| 2005 | An evaluation of evolved term-weighting schemes in information retrievalabstractThis paper presents an evaluation of evolved term-weighting schemes on short, medium and long TREC queries. A previously evolved global (collection-wide) term-weighting scheme is evaluated on unseen TREC data and is shown to increase mean average precision over idf. A local (within-document) evolved term-weighting scheme is presented which is dependent on the best performing global scheme. The full evolved scheme (i.e. the combined local and global scheme) is compared to both the BM25 scheme and the Pivoted Normalisation scheme.Our results show that the local evolved solution does not perform well on some collections due to its document normalisation properties and we conclude that Okapi-tf can be tuned to interact effectively with the evolved global weighting scheme presented and increase mean average precision over the standard BM25 scheme. Ronan Cummins, Colm O'Riordan |
CIKM | 1 |