Timothy Baldwin

dblp:65/4863 · also Tim Baldwin · DBLP profile ↗
← Back
15ranked-venue papers in the field
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 12Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Uncertainty Quantification for Large Language Models
Maxim Panov, Artem Shelmanov, Roman Vashurin, Artem Vazhentsev, Ekaterina Fadeeva, Lyudmila Rvanova, Timothy Baldwin
ECIR (4)7
2022 The ChEMU 2022 Evaluation Campaign: Information Extraction in Chemical Patents
Yuan Li 0012, Biaoyan Fang, Estrid He, Hiyori Yoshikawa, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zenan Zhai, Zubair Afzal, Trevor Cohn, Timothy Baldwin, Karin Verspoor
ECIR (2)11
2021 ChEMU 2021: Reaction Reference Resolution and Anaphora Resolution in Chemical Patents
Estrid He, Biaoyan Fang, Hiyori Yoshikawa, Yuan Li 0012, Saber A. Akhondi, Christian Druckenbrodt, Camilo Thorne, Zubair Afzal, Zenan Zhai, Lawrence Cavedon, Trevor Cohn, Timothy Baldwin, Karin Verspoor
ECIR (2)12
2021 Brief Description of COVID-SEE: The Scientific Evidence Explorer for COVID-19 Related Research
Karin Verspoor, Simon Suster, Yulia Otmakhova 0001, Shevon Mendis, Zenan Zhai, Biaoyan Fang, Jey Han Lau, Timothy Baldwin, Antonio Jimeno-Yepes, David Martínez 0001
ECIR (2)8
2020 ChEMU: Named Entity Recognition and Event Extraction of Chemical Reactions from Patents
Dat Quoc Nguyen, Zenan Zhai, Hiyori Yoshikawa, Biaoyan Fang, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Saber A. Akhondi, Trevor Cohn, Timothy Baldwin, Karin Verspoor
ECIR (2)10
2018 Detecting Misflagged Duplicate Questions in Community Question-Answering Archives
Doris Hoogeveen, Andrew Bennett, Yitong Li 0002, Karin Verspoor, Timothy Baldwin
ICWSM5
2018 A Living Lab Study of Query Amendment in Job Search
abstract
Errors in formulation of queries made by users can lead to poor search results pages. We performed a living lab study using online A/B testing to measure the degree of improvement achieved with a query amendment technique when applied to a commercial job search engine. Of particular interest in this case study is a clear 'success' signal, namely, the number of job applications lodged by a user as a result of querying the service. A set of 276 queries was identified for amendment in four different categories through the use of word embeddings, with large gains in conversion rates being attained in all four of those categories. Our analysis of query reformulations also provides a better understanding of user satisfaction in the case of problematic queries (ones with fewer results than fill a single page) by observing that users tend to reformulate rewritten queries less.
Bahar Salehi, Damiano Spina, Alistair Moffat, Seyedeh Sargol Sadeghi, Falk Scholer, Timothy Baldwin, Lawrence Cavedon, Mark Sanderson, Wilson Wong, Justin Zobel
SIGIR6
2018 The Company They Keep: Extracting Japanese Neologisms Using Language Patterns
abstract
We describe an investigation into the identification and extraction of unrecorded potential lexical items in Japanese text by detecting text passages containing selected language patterns typically associated with such items.We identified a set of suitable patterns, then tested them with two large collections of text drawn from the WWW and Twitter.Samples of the extracted items were evaluated, and it was demonstrated that the approach has considerable potential for identifying terms for later lexicographic analysis.
James Breen, Timothy Baldwin, Francis Bond
GWC2
2017 Evaluating topic representations for exploring document collections
abstract
Topic models have been shown to be a useful way of representing the content of large document collections, for example, via visualization interfaces (topic browsers). These systems enable users to explore collections by way of latent topics. A standard way to represent a topic is using a term list; that is the top‐n words with highest conditional probability within the topic. Other topic representations such as textual and image labels also have been proposed. However, there has been no comparison of these alternative representations. In this article, we compare 3 different topic representations in a document retrieval task. Participants were asked to retrieve relevant documents based on predefined queries within a fixed time limit, presenting topics in one of the following modalities: (a) lists of terms, (b) textual phrase labels, and (c) image labels. Results show that textual labels are easier for users to interpret than are term lists and image labels. Moreover, the precision of retrieved documents for textual and image labels is comparable to the precision achieved by representing topics using term lists, demonstrating that labeling methods are an effective alternative topic representation.
Nikolaos Aletras, Timothy Baldwin, Jey Han Lau, Mark Stevenson 0001
J. Assoc. Inf. Sci. Technol.2
2016 Quit While Ahead: Evaluating Truncated Rankings
abstract
Many types of search tasks are answered through the computation of a ranked list of suggested answers. We re-examine the usual assumption that answer lists should be as long as possible, and suggest that when the number of matching items is potentially small -- perhaps even zero -- it may be more helpful to "quit while ahead", that is, to truncate the answer ranking earlier rather than later. To capture this effect, metrics are required which are attuned to the length of the ranking, and can handle cases in which there are no relevant documents. In this work we explore a generalized approach for representing truncated result sets, and propose modifications to a number of popular evaluation metrics.
Fei Liu 0023, Alistair Moffat, Timothy Baldwin, Xiuzhen Zhang 0001
SIGIR3
2015 TM 2015 - Topic Models: Post-Processing and Applications Workshop
abstract
The main objective of the workshop is to bring together researchers who are interested in applications of topic models and improving their output. Our goal is to create a broad platform for researchers to share ideas that could improve the usability and interpretation of topic models. We expect this will promote topic model applications in other research areas, making their use more effective.
Nikolaos Aletras, Jey Han Lau, Timothy Baldwin, Mark Stevenson 0001
CIKM3
2015 A Probabilistic Rating Auto-encoder for Personalized Recommender Systems
abstract
User profiling is a key component of personalized recommender systems, and is used to generate user profiles that describe individual user interests and preferences. The increasing availability of big data is driving the urgent need for user profiling algorithms that are able to generate accurate user profiles from large-scale user behavior data. In this paper, we propose a probabilistic rating auto-encoder to perform unsupervised feature learning and generate latent user feature profiles from large-scale user rating data. Based on the generated user profiles, neighbourhood based collaborative filtering approaches have been adopted to make personalized rating predictions. The effectiveness of the proposed approach is demonstrated in experiments conducted on a real-world rating dataset from yelp.com.
Huizhi Liang 0001, Timothy Baldwin
CIKM2
2013 Lexical normalization for social media text
abstract
Twitter provides access to large volumes of data in real time, but is notoriously noisy, hampering its utility for NLP. In this article, we target out-of-vocabulary words in short text messages and propose a method for identifying and normalizing lexical variants. Our method uses a classifier to detect lexical variants, and generates correction candidates based on morphophonemic similarity. Both word similarity and context are then exploited to select the most probable correction candidate for the word. The proposed method doesn't require any annotations, and achieves state-of-the-art performance over an SMS corpus and a novel dataset based on Twitter.
Bo Han 0002, Paul Cook, Timothy Baldwin
ACM Trans. Intell. Syst. Technol.3
2010 Visualizing search results and document collections using topic maps
David Newman 0001, Timothy Baldwin, Lawrence Cavedon, Sarvnaz Karimi, David Martínez 0001, Falk Scholer, Justin Zobel
J. Web Semant.2
2009 Experiments on pattern-based relation learning
abstract
Relation extraction is the task of extracting semantic relations - such as synonymy or hypernymy - between word pairs from corpus data. Past work in relation extraction has concentrated on manually creating templates to use in directly extracting word pairs for a given semantic relation from corpus text. Recently, there has been a move towards using machine learning to automatically learn these patterns. We build on this research by running experiments investigating the impact of corpus type, corpus size and different parameter settings on learning a range of lexical relations.
Willy Yap, Timothy Baldwin
CIKM2