VLDB 2026 Research / reviewers in the wild / expert
Hideo Joho
dblp:83/5350
· DBLP profile ↗
48ranked-venue papers in the field
14as first author
7since 2021 · last 2026
0000-0002-6611-652XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 42 (14 first)Database Systems & Data Management · 4Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CollabSearch: A Study of User-LLM Collaboration in Task-Based SearchabstractWith the rapid proliferation of large language models (LLMs), users are increasingly turning to these systems to fulfill their everyday information needs. Unlike traditional search engines, which rely on structured query-response mechanisms, LLMs offer direct answer synthesis—fundamentally altering how users access and interact with information. However, the effectiveness of such direct responses and how users interact with such answers were not studied. In this work, we present CollabSearch, a user-LLM collaborative search system that enables users to interact with LLMs through adaptive role-playing prompts designed to guide and refine search outcomes. Users leverage the internal knowledge of LLMs while also benefiting from retrieval-augmented generation (RAG) to ensure relevance and currency. We evaluate our system through a task-based study involving 24 participants. The results demonstrate the effectiveness of our approach and offer nuanced insights into user–LLM interaction dynamics in search contexts. Gloris Denisse Cedeño Batista, Joemon M. Jose, Hideo Joho |
CHIIR | 3 |
| 2026 | Controlled Experimentation of Model Search Behaviour with Geniie-LababstractThis half-day tutorial provides participants with a framework and hands-on experience for conducting controlled experiments on model search behaviour using an open-source toolkit. Participants will learn how to design, run, and analyse experiments to investigate behavioural differences across large language models. By integrating techniques from human user studies with LLM experimentation, the tutorial strengthens CHIIR’s methodological foundations and broadens its scope to include behavioural analysis of generative and agentic systems. Hideo Joho |
CHIIR | 1 |
| 2026 | Retrieval-Augmented Diffusion Language Model for Generative Commonsense Reasoning
Yubo Fang, Hai-Tao Yu 0003, Hideo Joho, Sumio Fujita |
DASFAA (3) | 3 |
| 2025 | Eliciting Implicit Information Needs in E-Commerce Search by Using the Think-Aloud MethodabstractIn this paper, we elicit implicit information needs that arise during the process of deciding which products to purchase on e-commerce (EC) sites.We designed product purchase tasks to capture implicit information needs, and we conducted a user study to collect utterance data using a think-aloud method.By analyzing the utterances of participants during the tasks, we developed a taxonomy comprising five categories where people express preferences for products and 11 categories where people want to understand products.Our taxonomy includes implicit information needs that have not been captured in existing EC-related taxonomies (e.g., Preference for Subjective Attributes and Understanding Product Differences).We revealed the characteristics of each category of information need in terms of timing during the tasks: e.g., the information need of Understanding Product Range occurred very frequently in the early stage of a task.We also revealed the occurrence frequencies for different task types: e.g., the information needs of Preference for Objective Attributes, Understanding Product Range, and Understanding Terminology had a higher occurrence when purchasing products less frequently and at a higher cost than when purchasing products frequently at a relatively low cost.Our taxonomy could be used to further improve users' purchasing processes on EC sites. Kosetsu Tsukuda, Atsuki Maruta, Makoto P. Kato, Hideo Joho |
CHIIR | 4 |
| 2025 | An Instruction-Response Perspective on Large Language Models in Information Retrieval TasksabstractThe increasing use of retrieval-augmented applications, where large language models (LLMs) are instructed to generate queries, assess relevance, and synthesise responses, has introduced new challenges in Information Retrieval (IR).The lack of transparency in LLMs means that even subtle variations in instructions can significantly impact the quality, consistency, and reliability of their responses.To address this issue, we propose Instruction-Response Study, an experimental framework for systematically analysing how task instructions influence LLM-generated responses in IR tasks.This paper presents the core components of the framework and demonstrates its utility through four case studies, examining 1) the effect of IR tasks on query formulation, 2) the impact of topic information size on retrieval effectiveness, 3) the reproducibility of LLM-generated queries, and 4) the role of meta-instructions in diversifying instruction design.The findings highlight how the proposed framework enables controlled experimentation on instruction design and its effects, offering a foundation for optimising prompt engineering and enhancing retrieval-augmented applications. Hideo Joho, Joemon M. Jose |
SIGIR | 1 |
| 2023 | An in-depth study on adversarial learning-to-rank
Hai-Tao Yu 0003, Rajesh Piryani, Adam Jatowt, Ryo Inagaki, Hideo Joho, Kyoung-Sook Kim 0001 |
Inf. Retr. J. | 5 |
| 2022 | Selectively Expanding Queries and Documents for News Background LinkingabstractBackground articles are crucial for readers to grasp the context of news stories fully. However, existing approaches of background article search tend to apply a single ranking method to all types of search topics. In this paper, we focus on exploring search topics on news articles by classifying them into two types:time-sensitive andnon-time-sensitive. To verify whether or not these two types of search topics can benefit from different retrieving methods, we examined a suite of strategies such as document expansion, query rewriting, and semantic re-ranking. Moreover, the relationship between background articles and topics is verified by the two strategies of document expansion (specificity and diversity). The experimental results demonstrate that the optimal usage of the aforementioned strategies is indeed different between the two types of search topics. Furthermore, our in-depth analysis of topics and search results verified that: time-sensitive topics benefit from background articles that can provide more specific knowledge, while non-time-sensitive topics benefit from diversified retrieved documents. Lirong Zhang, Hideo Joho, Sumio Fujita, Hai-Tao Yu 0003 |
CIKM | 2 |
| 2020 | Third International Workshop on Conversational Approaches to Information Retrieval (CAIR'20): Full-day Workshop at CHIIR 2020abstractThe third CAIR workshop brings together researchers and developers interested in advancing conversational systems in interactive information retrieval. The workshop builds on the first and second CAIR workshops held at SIGIR 2017 and 2018 and will focus on the continuing development of current challenges, user and system limitations, and evaluation of conversational systems for information retrieval. Participants will collaboratively explore different contexts (i.e., home, hospitals, or work settings), use cases, and interactivity forms (voice-only, multi-modal, screen-based) in which conversational search systems can be used. Possible outcomes include fostering novel and innovative methodologies (such as for data collection and evaluation), personalising conversational systems, and understanding ethical challenges---such as system transparency---from the user's perspective. Johanne R. Trippas, Paul Thomas 0001, Damiano Spina, Hideo Joho |
CHIIR | 4 |
| 2020 | Towards a model for spoken conversational search
Johanne R. Trippas, Damiano Spina, Paul Thomas 0001, Mark Sanderson, Hideo Joho, Lawrence Cavedon |
Inf. Process. Manag. | 5 |
| 2020 | A Price-per-attention Auction Scheme Using Mouse Cursor InformationabstractPayments in online ad auctions are typically derived from click-through rates, so that advertisers do not pay for ineffective ads. But advertisers often care about more than just clicks. That is, for example, if they aim to raise brand awareness or visibility. There is thus an opportunity to devise a more effective ad pricing paradigm, in which ads are paid only if they are actually noticed. This article contributes a novel auction format based on a pay-per-attention (PPA) scheme. We show that the PPA auction inherits the desirable properties (strategy-proofness and efficiency) as its pay-per-impression and pay-per-click counterparts, and that it also compares favourably in terms of revenues. To make the PPA format feasible, we also contribute a scalable diagnostic technology to predict user attention to ads in sponsored search using raw mouse cursor coordinates only, regardless of the page content and structure. We use the user attention predictions in numerical simulations to evaluate the PPA auction scheme. Our results show that, in relevant economic settings, the PPA revenues would be strictly higher than the existing auction payment schemes. Ioannis Arapakis, Antonio Penta, Hideo Joho, Luis A. Leiva |
ACM Trans. Inf. Syst. | 3 |
| 2019 | WassRank: Listwise Document Ranking Using Optimal Transport TheoryabstractLearning to rank has been intensively studied and has shown great value in many fields, such as web search, question answering and recommender systems. This paper focuses on listwise document ranking, where all documents associated with the same query in the training data are used as the input. We propose a novel ranking method, referred to as WassRank, under which the problem of listwise document ranking boils down to the task of learning the optimal ranking function that achieves the minimum Wasserstein distance. Specifically, given the query level predictions and the ground truth labels, we first map them into two probability vectors. Analogous to the optimal transport problem, we view each probability vector as a pile of relevance mass with peaks indicating higher relevance. The listwise ranking loss is formulated as the minimum cost (the Wasserstein distance) of transporting (or reshaping) the pile of predicted relevance mass so that it matches the pile of ground-truth relevance mass. The smaller the Wasserstein distance is, the closer the prediction gets to the ground-truth. To better capture the inherent relevance-based order information among documents with different relevance labels and lower the variance of predictions for documents with the same relevance label, ranking-specific cost matrix is imposed. To validate the effectiveness of WassRank, we conduct a series of experiments on two benchmark collections. The experimental results demonstrate that: compared with four non-trivial listwise ranking methods (i.e., LambdaRank, ListNet, ListMLE and ApxNDCG), WassRank can achieve substantially improved performance in terms of nDCG and ERR across different rank positions. Specifically, the maximum improvements of WassRank over LambdaRank, ListNet, ListMLE and ApxNDCG in terms of [email protected] are 15%, 5%, 7%, 5%, respectively. Hai-Tao Yu 0003, Adam Jatowt, Hideo Joho, Joemon M. Jose, Long Chen 0008 |
WSDM | 3 |
| 2019 | An analysis of natural disaster-related information-seeking behavior using temporal stagesabstractSince natural disasters can affect many people over a vast area, studying information‐seeking behavior (ISB) during disasters is of great importance. Many previous studies have relied on online social network data, providing insights into the ISB of those with Internet access. However, in a large‐scale natural disaster such as the Great East Japan Earthquake of 2011, people in the most severely affected areas tended to have limited Internet access. Therefore, an alternative data source should be explored to investigate disaster‐related ISB. This study's contributions are twofold. First, we provide a detailed description of natural disaster‐related ISB of people who experienced a large‐scale earthquake and tsunami, based on analysis of written testimonies published by local authorities. This provided insight into the relationship between information needs, channels, and sources of disaster‐related ISB. Also, our approach facilitates the study of ISB of people without Internet access both during and after a disaster. Second, we provide empirical evidence to demonstrate that the temporal stages of a disaster can characterize people's ISB during the disaster. Therefore, we propose further consideration of the temporal aspects of events for improved understanding of disaster‐related ISB. Rahmi Rahmi, Hideo Joho, Tetsuya Shirai |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2018 | Informing the Design of Spoken Conversational Search: Perspective PaperabstractWe conducted a laboratory-based observational study where pairs of people performed search tasks communicating verbally. Examination of the discourse allowed commonly used interactions to be identified for Spoken Conversational Search (SCS). We compared the interactions to existing models of search behaviour. We find that SCS is more complex and interactive than traditional search. This work enhances our understanding of different search behaviours and proposes research opportunities for an audio-only search system. Future work will focus on creating models of search behaviour for SCS and evaluating these against actual SCS systems. Johanne R. Trippas, Damiano Spina, Lawrence Cavedon, Hideo Joho, Mark Sanderson |
CHIIR | 4 |
| 2018 | Second International Workshop on Conversational Approaches to Information Retrieval (CAIR'18): Workshop at SIGIR 2018abstractThe CAIR'18 workshop will bring together academic and industrial researchers to create a forum for research on conversational approaches to search and recommendation. A specific focus will be on techniques that support complex and multi-turn user-machine dialogues for information access and retrieval, and multi-modal interfaces for interacting with such systems. Jaime Arguello, Filip Radlinski, Hideo Joho, Damiano Spina, Julia Kiseleva |
SIGIR | 3 |
| 2018 | Revisiting the cluster-based paradigm for implicit search result diversification
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose, Long Chen 0008, Fajie Yuan |
Inf. Process. Manag. | 4 |
| 2017 | First International Workshop on Conversational Approaches to Information Retrieval (CAIR'17)abstractRecent advances in commercial conversational services that allow naturally spoken and typed interaction, particularly for well-formulated questions and commands, have increased the need for more human-centric interactions in information retrieval. The First International Workshop on Conversational Approaches to Information Retrieval (CAIR`17) brings together academic and industrial researchers to create a forum for research on conversational approaches to search. A specific focus is on techniques that support complex and multi-turn user-machine dialogues for information access and retrieval, and multi-model interfaces for interacting with such systems. We invite submissions addressing all modalities of conversation, including speech-based, text-based, and multimodal interaction. We also welcome studies of human-human interaction (e.g., collaborative search) that can inform the design of conversational search applications, and work on evaluation of conversational approaches. Hideo Joho, Lawrence Cavedon, Jaime Arguello, Milad Shokouhi, Filip Radlinski |
SIGIR | 1 |
| 2017 | Modelling Information Needs in Collaborative Search ConversationsabstractThe increase of voice-based interaction has changed the way people seek information, making search more conversational. Development of effective conversational approaches to search requires better understanding of how people express information needs in dialogue. This paper describes the creation and examination of over 32K spoken utterances collected during 34 hours of collaborative search tasks. The contribution of this work is three-fold. First, we propose a model of conversational information needs (CINs) based on a synthesis of relevant theories in Information Seeking and Retrieval. Second, we show several behavioural patterns of CINs based on the proposed model. Third, we identify effective feature groups that may be useful for detecting CINs categories from conversations. This paper concludes with a discussion of how these findings can facilitate advance of conversational search applications. Sosuke Shiga, Hideo Joho, Roi Blanco, Johanne R. Trippas, Mark Sanderson |
SIGIR | 2 |
| 2017 | A Concise Integer Linear Programming Formulation for Implicit Search Result DiversificationabstractTo cope with ambiguous and/or underspecified queries, search result diversification (SRD) is a key technique that has attracted a lot of attention. This paper focuses on implicit SRD, where the possible subtopics underlying a query are unknown beforehand. We formulate implicit SRD as a process of selecting and ranking k exemplar documents that utilizes integer linear programming (ILP). Unlike the common practice of relying on approximate methods, this formulation enables us to obtain the optimal solution of the objective function. Based on four benchmark collections, our extensive empirical experiments reveal that: (1) The factors, such as different initial runs, the number of input documents, query types and the ways of computing document similarity significantly affect the performance of diversification models. Careful examinations of these factors are highly recommended in the development of implicit SRD methods. (2) The proposed method can achieve substantially improved performance over the state-of-the-art unsupervised methods for implicit SRD. Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose, Long Chen 0008, Fajie Yuan |
WSDM | 4 |
| 2017 | An in-depth study on diversity evaluation: The importance of intrinsic diversity
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose |
Inf. Process. Manag. | 4 |
| 2017 | Decoding multi-click search behavior based on marginal utility
Hai-Tao Yu 0003, Adam Jatowt, Roi Blanco, Hideo Joho, Joemon M. Jose |
Inf. Retr. J. | 4 |
| 2016 | System And User Centered Evaluation Approaches in Interactive Information Retrieval (SAUCE 2016)abstractThe purpose of this half-day workshop is to bring together academic and industry interactive information retrieval (IIR) researchers with an interest in evaluation methodologies. The workshop articulates contemporary challenges in the investigation of IIR and invites user- and system-oriented researchers to work collaboratively to address these challenges by combining user- and system-centered methodologies in meaningful ways. We anticipate that this workshop will initiate productive knowledge exchange and partnerships that can respond to the increasing user, task, system, and contextual complexity of the IIR field. Heather L. O'Brien, Nicola Ferro 0001, Hideo Joho, Dirk Lewandowski, Paul Thomas 0001, C. J. van Rijsbergen |
CHIIR | 3 |
| 2016 | NTCIR Lifelog: The First Test Collection for Lifelog ResearchabstractTest collections have a long history of supporting repeatable and comparable evaluation in Information Retrieval (IR). However, thus far, no shared test collection exists for IR systems that are designed to index and retrieve multimodal lifelog data. In this paper we introduce the first test collection for personal lifelog data, which has been employed for the NTCIR12-Lifelog task. In this paper, the requirements for the test collection are motivated, the process of creating the test collection is described, along with an overview of the test collection. Finally suggestions are given for possible applications of the test collection. Cathal Gurrin, Hideo Joho, Frank Hopfgartner, Liting Zhou, Rami Albatal |
SIGIR | 2 |
| 2016 | Building Test Collections for Evaluating Temporal IRabstractResearch on temporal aspects of information retrieval has recently gained considerable interest within the Information Retrieval (IR) community. This paper describes our efforts for building test collections for the purpose of fostering temporal IR research. In particular, we overview the test collections created at the two recent editions of Temporal Information Access (Temporalia) task organized at NTCIR-11 and NTCIR-12, report on selected results and discuss several observations we made during the task design and implementation. Finally, we outline further directions for constructing test collections suitable for temporal IR. Hideo Joho, Adam Jatowt, Roi Blanco, Hai-Tao Yu 0003, Shuhei Yamamoto |
SIGIR | 1 |
| 2015 | Temporal information searching behaviour and strategies
Hideo Joho, Adam Jatowt, Roi Blanco |
Inf. Process. Manag. | 1 |
| 2013 | Tempo of Search Actions to Modeling Successful Sessions
Kazuya Fujikawa, Hideo Joho, Shin-ichi Nakayama |
ECIR | 2 |
| 2013 | Doctoral Consortium at ECIR 2013
Hideo Joho, Dmitry I. Ignatov |
ECIR | 1 |
| 2011 | Formulating effective questions for community-based question answeringabstractCommunity-based Question Answering (CQA) services have become a major venue for people's information seeking on the Web. However, many studies on CQA have focused on the prediction of the best answers for a given question. This paper looks into the formulation of effective questions in the context of CQA. In particular, we looked at effect of contextual factors appended to a basic question on the performance of submitted answers. This study analysed a total of 930 answers returned in response to 266 questions that were formulated by 46 participants. The results show that adding a questionnaire's personal and social attribute to the question helped improve the perceptions of answers both in information seeking questions and opinion seeking questions. Saori Suzuki, Shin-ichi Nakayama, Hideo Joho |
SIGIR | 3 |
| 2011 | Study of context influence on classifiers trained under different video-document representations
Pablo Bermejo 0001, Hideo Joho, Joemon M. Jose, Robert Villa |
Inf. Process. Manag. | 2 |
| 2011 | Diane Kelly: Methods for evaluating interactive information retrieval systems with users - Foundation and Trends in Information Retrieval, vol 3, nos 1-2, pp 1-224, 2009, ISBN: 978-1-60198-224-7
Hideo Joho |
Inf. Retr. | 1 |
| 2010 | Factors affecting click-through behavior in aggregated search interfacesabstractAn aggregated search interface is designed to integrate search results from different sources (web, image, video, blog, etc) into a single result page. This paper presents two user studies investigating factors affecting users click-through behavior on aggregated search interfaces. We tested two aggregated search interfaces: one where results from the different sources are blended into a single list (called blended), and another, where results from each source are presented in a separate panel (called non-blended). A total of 1,296 search sessions performed by 48 participants were analysed in our study. Our results suggest that 1) the position of search results is significant only in the blended and not in the non-blended design; 2) participants' click-through behavior on videos is different from other sources; and finally 3) capturing a task's orientation towards particular sources is an important factor for further investigation and research. Shanu Sushmita, Hideo Joho, Mounia Lalmas-Roelleke, Robert Villa |
CIKM | 2 |
| 2010 | Search system requirements of patent analystsabstractPatent search tasks are difficult and challenging, often requiring expert patent analysts to spend hours, even days, sourcing relevant information. To aid them in this process, analysts use Information Retrieval systems and tools to cope with their retrieval tasks. With the growing interest in patent search, it is important to determine their requirements and expectations of the tools and systems that they employ. In this poster, we report a subset of the findings of a survey of patent analysts conducted to elicit their search requirements. Leif Azzopardi, Wim Vanderbauwhede, Hideo Joho |
SIGIR | 3 |
| 2009 | Revisiting IR Techniques for Collaborative Search Strategies
Hideo Joho, David Hannah, Joemon M. Jose |
ECIR | 1 |
| 2009 | An aspectual interface for supporting complex search tasksabstractWith the increasing importance of search systems on the web, there is a continuing push to design interfaces which are a better match with the kinds of real-world tasks in which users are engaged. In this paper, we consider how broad, complex search tasks may be supported via the search interface. In particular, we consider search tasks which may be composed of multiple aspects, or multiple related subtasks. For example, in decision making tasks the user may investigate multiple possible solutions before settling on a single, final solution, while other tasks, such as report writing, may involve searching on multiple interrelated topics. Robert Villa, Iván Cantador, Hideo Joho, Joemon M. Jose |
SIGIR | 3 |
| 2009 | A Task-Based Evaluation of an Aggregated Search Interface
Shanu Sushmita, Hideo Joho, Mounia Lalmas-Roelleke |
SPIRE | 2 |
| 2008 | Emulating query-biased summaries using document titlesabstractGenerating query-biased summaries can take up a large part of the response time of interactive information retrieval (IIR) systems. This paper proposes to use document titles as an alternative to queries in the generation of summaries. The use of document titles allows us to pre-generate summaries statically, and thus, improve the response speed of IIR systems. Our experiments suggest that title-biased summaries are a promising alternative to query-biased summaries. Hideo Joho, David Hannah, Joemon M. Jose |
SIGIR | 1 |
| 2008 | Modelling vague places with knowledge from the WebabstractPlace names are often used to describe and to enquire about geographical information. It is common for users to employ vernacular names that have vague spatial extent and which do not correspond to the official and administrative place name terminology recorded within typical gazetteers. There is a need therefore to enrich gazetteers with knowledge of such vague places and hence improve the quality of place name‐based information retrieval. Here we describe a method for modelling vague places using knowledge harvested from Web pages. It is found that vague place names are frequently accompanied in text by the names of more precise co‐located places that lie within the extent of the target vague place. Density surface modelling of the frequency of co‐occurrence of such names provides an effective method of representing the inherent uncertainty of the extent of the vague place while also enabling approximate crisp boundaries to be derived from contours if required. The method is evaluated using both precise and vague places. The use of the resulting approximate boundaries is demonstrated using an experimental geographical search engine. Christopher B. Jones, Ross Purves, Paul D. Clough, Hideo Joho |
Int. J. Geogr. Inf. Sci. | 4 |
| 2008 | Effectiveness of additional representations for the search result presentation on the web
Hideo Joho, Joemon M. Jose |
Inf. Process. Manag. | 1 |
| 2008 | Adaptive information retrieval: Introduction to the special topic issue of information processing and management
Joemon M. Jose, Hideo Joho, C. J. van Rijsbergen |
Inf. Process. Manag. | 2 |
| 2007 | Evaluating Query-Independent Object Features for Relevancy Prediction
Andrés R. Masegosa, Hideo Joho, Joemon M. Jose |
ECIR | 2 |
| 2007 | Effects of highly agreed documents in relevancy predictionabstractFinding significant contextual features is a challenging task in the development of interactive information retrieval (IR) systems. This paper investigated a simple method to facilitate such a task by looking at aggregated relevance judgements of retrieved documents. Our study suggested that the agreement on relevance judgements can indicate the effectiveness of retrieved documents as the source of significant features. The effect of highly agreed documents gives us practical implication for the design of adaptive search models in interactive IR systems. Andrés R. Masegosa, Hideo Joho, Joemon M. Jose |
SIGIR | 2 |
| 2007 | The design and implementation of SPIRIT: a spatially aware search engine for information retrieval on the InternetabstractMuch of the information stored on the web contains geographical context, but current search engines treat such context in the same way as all other content. In this paper we describe the design, implementation and evaluation of a spatially aware search engine which is capable of handling queries in the form of the triplet of ⟨theme⟩⟨spatial relationship⟩⟨location⟩. The process of identifying geographic references in documents and assigning appropriate footprints to documents, to be stored together with document terms in an appropriate indexing structure allowing real‐time search, is described. Methods allowing users to query and explore results which have been relevance‐ranked in terms of both thematic and spatial relevance have been implanted and a usability study indicates that users are happy with the range of spatial relationships available and intuitively understand how to use such a search engine. Normalised precision for 38 queries, containing four types of spatial relationships, is significantly higher (p<0.001) for searches exploiting spatial information than pure text search. Ross Purves, Paul D. Clough, Christopher B. Jones, Avi Arampatzis, Bénédicte Bucher, David Finch, Gaihua Fu, Hideo Joho, Awase Khirni Syed, Subodh Vaid, Bisheng Yang |
Int. J. Geogr. Inf. Sci. | 8 |
| 2006 | Judging the Spatial Relevance of Documents for GIR
Paul D. Clough, Hideo Joho, Ross Purves |
ECIR | 2 |
| 2006 | A Comparative Study of the Effectiveness of Search Result Presentation on the Web
Hideo Joho, Joemon M. Jose |
ECIR | 1 |
| 2005 | Spatio-textual Indexing for Geographical Search on the Web
Subodh Vaid, Christopher B. Jones, Hideo Joho, Mark Sanderson |
SSTD | 3 |
| 2004 | A Study of User Interaction with a Concept-Based Interactive Query Expansion Support Tool
Hideo Joho, Mark Sanderson, Micheline Beaulieu |
ECIR | 1 |
| 2004 | Forming test collections with no system poolingabstractForming test collection relevance judgments from the pooled output of multiple retrieval systems has become the standard process for creating resources such as the TREC, CLEF, and NTCIR test collections. This paper presents a series of experiments examining three different ways of building test collections where no system pooling is used. First, a collection formation technique combining manual feedback and multiple systems is adapted to work with a single retrieval system. Second, an existing method based on pooling the output of multiple manual searches is re-examined: testing a wider range of searchers and retrieval systems than has been examined before. Third, a new approach is explored where the ranked output of a single automatic search on a single retrieval system is assessed for relevance: no pooling whatsoever. Using established techniques for evaluating the quality of relevance judgments, in all three cases, test collections are formed that are as good as TREC. Mark Sanderson, Hideo Joho |
SIGIR | 2 |
| 2002 | Hierarchical approach to term suggestion deviceabstractOur demonstration shows the hierarchy system working on a locally run search engine. Hierarchies are dynamically generated from the retrieved documents, and visualised on the menus. When a user selects a term from the hierarchy, the documents linked to the term are listed, and the term is then added to the initial query to rerun a search. Through the demonstration we illustrate how hierarchical presentation of expansion terms is achieved, and how our approach supports users to articulate their information needs using the hierarchy. Hideo Joho, Mark Sanderson, Micheline Beaulieu |
SIGIR | 1 |
| 2000 | Retrieving Descriptive Phrases from Large Amounts of Free TextabstractThis paper presents a system that retrieves descriptive phrases of proper nouns from free text. Sentences holding the specified noun are ranked using a technique based on pattern matching, word counting, and sentence location. No domain specific knowledge is used. Experiments show the system able to rank highly those sentences that contain phrases describing or defining the query noun. In contrast to existing methods, this system does not use parsing techniques but still achieves high levels of accuracy. From the results of a large-scale experiment, it is speculated that the success of this simpler method is due to the high quantities of free text being searched. Parallels between this work and recent findings in the very large corpus track of TREC are drawn. Keywords Information retrieval, descriptive phrase, large corpora. 1. INTRODUCTION The opportunities to use online text databases for the mining of valuable information are great. As these stores increase in size, the possibil... Hideo Joho, Mark Sanderson |
CIKM | 1 |